Logo image
A study of enzyme class prediction - from functional domain composition and phylogenetic relationships of protein sequence
Dissertation

A study of enzyme class prediction - from functional domain composition and phylogenetic relationships of protein sequence

Shih-Hau Chiu
Doctor of Philosophy (PHD), 國立清華大學, 分子醫學研究所
2007

Abstract

關聯演算法 支持向量機演算法 功能區域組成 物理化學特性 類源分類器 類源組別 association algorithm Apriori support vector machines functional domain composition enzyme class InterPro entries physicochemical features phylogenetics
Identifying the function of protein sequence is still a challenging task in computational and informational biology. Traditionally, functional prediction mostly relies on detecting similarity between a functionally annotated protein and the query protein, then transferring the annotations across. However, sequence composition bias influence the results of similarity searches, and they do not yield the exact share between biological function and domain composition based on the similarity threshold used. Moreover, for the enzyme class prediction, the predictive capacity of previous studies are just to the top level or sublevel of EC classification system, which no research findings are yet available concerning the exact (four-digit) EC numbers prediction. In this work, we attempt to construct a work flow for automatic mapping protein sequence to their corresponding enzyme class based on functional domain composition of protein. The association algorithm, Apriori, is utilized to mine the relationship between the enzyme class and significant InterPro entries. The candidate rules are evaluated for their classificatory capacity. A correct enzyme classification rate of 70% was obtained for the prokaryote datasets and a similar rate of about 80% was obtained for the eukaryote datasets. Furthermore, we found that the rules were different among five taxonomic datasets studied. Consequently, to use these rules, one has to know the phylogeny of protein sequences beforehand. Here, we provide a straightforward method to predict the phylogeny of protein sequences by using a support vector machines classifier based on the biochemical features of amino acid sequences of the genomes. The classification accuracies of the trained SVM classifiers by the Enzymatic and All proteins are 84 and 79%, respectively. Results show that some compositions or biochemical features of amino acid sequences of the genomes can be used to cluster proteins of different taxonomic natures. The sequence compositions of proteins analyzed are originated from some special characteristics corresponding to the taxonomic clades. We prove that the phylogenetic class of protein sequence can be predicted just by amino acid physicochemical properties alone.

Metrics

1 Record Views

Details

Logo image