Logo image
Coding Exon Prediction Based on Phylogenetical Comparisons
Dissertation

Coding Exon Prediction Based on Phylogenetical Comparisons

Shu-Ju Hsieh
Doctor of Philosophy (PHD), 國立清華大學, 資訊工程學系
2005

Abstract

序列比對 比較基因體
Identifying protein coding genes is a challenging task in computational biology. With the rapid accumulation of genomic sequences for various organisms, it is now feasible to identify novel genes and exons by genomic comparisons. Sequence analysis is the base of bioinformatics, in which the sequence alignment is a major and fundamental task. Comparing coding regions among organisms can pinpoint functionally important parts of proteins, which are more conserved than the other parts bearing no functional significance. However, most of the currently available alignment programs, though providing optimal or near optimal results, have been limited by their computation speed. In this thesis, we propose a new method for coding region alignments. Based on a probabilistic filtration approach, CORAL (COding Region ALignment) is a linear time alignment tool. Integrating CORAL and signal detectors, we developed two programs, EXONALIGN and GeneAlign for coding exon prediction. EXONALIGN simultaneously aligns and predicts exons between homologous genes/ syntenic regions. To reduce computation time and improve prediction accuracy, EXONALIGN calculates strengths of intrinsic splice signals and applies CORAL to measure sequence homologies between regions flanked by pairs of candidate splice acceptor and donor of homologous genes. The performance of EXONALIGN was evaluated on the ROSETTA and the Projector data sets. The predictions obtained by EXONALIGN are comparable with those obtained by widely used gene prediction programs, confirming the benefit of importing the conservation of exon-intron structures into an exon/gene prediction tool. Finally, EXONALIGN was employed to explore novel human exons within the annotated human-mouse homologous genes. More than one hundred novel human and mouse exon pairs were predicted within annotated genes. These putative human exons are longer than 100 bp and show greater than 70% sequence conservation to the corresponding mouse exons. The KA/KS ratios of 75% of the predicted exons are smaller than 1, further supporting the likelihood that the majority of newly predicted human exons code for proteins. Furthermore, with increasing numbers of gene annotations verified by experiments, it is feasible to identify genes in the newly sequenced genomes by comparing to the annotated genes of phylogenetically close organisms. GeneAlign predicts protein coding genes by measuring the homologies between the sequence of a newly sequenced genome and the homologue of a related genome. GeneAlign was tested on Projector data set of 491 human-mouse homologous sequence pairs. At the gene level, both the average sensitivity and the average specificity of GeneAlign are 81%, and they are larger than 96% at the exon level. The rates of missing exons and wrong exons are smaller than 1%.

Metrics

1 Record Views

Details

Logo image