Logo image
使用支持向量機以n個數擷取編碼蛋白質及表決法對蛋白質細部分類的研究
Thesis

使用支持向量機以n個數擷取編碼蛋白質及表決法對蛋白質細部分類的研究

游景盛
Masters, 國立清華大學, 生命科學系
2001

Abstract

支持向量機 SVM SCOP n-peptide
Fold assignment directly from sequences is valuable in the prediction of protein structures. Unlike secondary structure prediction, where a local coding scheme of sequence information will usually suffice, fold identification calls for global protein descriptors as well local descriptors for the whole protein sequences. Previous studies have shown that machine learning methods can yield reasonable prediction accuracy of fold assignment directly from sequences by a variety of global sequence coding schemes. In this thesis, using global protein descriptors based on -peptide distribution, we apply the support vector machine method (SVM) to the 27 most populated folds that contain 386 representative proteins in the Structural Classification of Protein (SCOP) database. Our approach achieved a prediction accuracy 69.6% on an independent set, and 55.5% in the ten-fold cross validation, both of which are an order of magnitude higher than the current methods. Our results show that SVM using suitable global sequence coding schemes can significantly improve prediction in fold recognition from sequences, and should offer a useful tool in structure modeling.

Metrics

1 Record Views

Details

Logo image