Abstract
Currently, one of the important issues in bioinformatics is the prediction of novel genes in human genome. Genes with specifc structures are the targets for annotation in the three billions base-pairs of the human genome. Polyadenylation site, a structure at the terminus of a gene, involves a precise endonucleolytic cleavage of the pre-mRNA followed by synthesis of the polyA tail which is found at the 3' end of nearly every mature eukaryotic mRNA. The recognition of polyadenylation site is governed by at least two signals : One is 10-30 nucleotides upstream to the cleavage/polyadenylation site and named as polyA signal (PAS), a highly conserved hexamer AAUAAA (and the common variant AUUAAA). The other is 20-40 nucleotides downstream to the cleavage/polyadenylation site, the downstream element (DE) consisting of a much less well-characterized U or G-U rich sequence. In this thesis, we will provide a program for the prediction of human polyadenylation site by the detection of the PAS signal and the DE signal with dependency graphs and their expanded Bayesian networks. Then we will compare the accuracy of prediction with famous programs POLYAH and ERPIN, and show that our program performs the best results in the polyadenylation dataset of GeneBank.