Logo image
人類與巨猿間之遺傳距離
Thesis

人類與巨猿間之遺傳距離

陳豐奇
Masters, National Tsing Hua University
2001

Abstract

巨猿基因組遺傳距離有效族群分子鐘非編碼基因序列大規模排序重複序列 the great apesgenomic divergenceeffective population sizemolecular clocknoncoding sequencelarge-scale alignmentrepetitive sequence
This study focuses on two subjects: the divergences between human and the great apes, and an in-depth research of the divergence between human and chimpanzee using a relatively large data set.This study differs from previous ones in that it demonstrates an approach to efficiently obtain authentic genomic divergences without sacrificing accuracy or reliability. Most of the previous studies are flawed in that they include only small numbers of loci. Therefore, the results of the previous studies are susceptible to strong sampling bias. In contrast, the method employed in this study not only eliminates sampling bias but also proves to be highly efficient. Moreover, this study demonstrates an algorithm for large-scale sequence alignments that involve repetitive sequences. This pioneer research has significant impacts on future studies in molecular evolution.To study the genomic divergences among hominoids and to estimate the effective population size of the common ancestor of human and chimpanzee we selected 53 autosomal intergenic noncoding DNA segments from the human genome and sequenced them in a human, chimpanzee, gorilla, and orangutan. The average sequence divergence was only 1.24% ± 0.07% for the human-chimpanzee pair, 1.62% ± 0.08% for the human-gorilla pair, and 1.63% ± 0.08% for the chimpanzee-gorilla pair. These estimates, which were confirmed by additional data from GenBank, are substantially lower than previous ones, which included repetitive sequences and might have been based on less accurate sequence data. The average sequence divergences between orangutan and human, chimpanzee, and gorilla were 3.08% ± 0.11%, 3.12% ± 0.11% and 3.09% ± 0.11%, respectively, which are also substantially lower than previous estimates. The sequence divergences in other regions between hominoids were estimated from extensive data in GenBank and the literature, and Alus showed the highest divergence, followed in order by Y-linked noncoding regions, pseudogenes, synonymous sites, autosomal intergenic regions, X-linked noncoding regions, introns, and nonsynonymous sites. The neighbor-joining tree derived from the concatenated sequence of the 53 segments, 24,234 bp in length, supports the Homo-Pan clade with a 100% bootstrap value. However, when each segment is analyzed separately 22 of the 53 segments (~42%) give a tree that is incongruent with the species tree, suggesting a large effective population size (Ne) of the common ancestor of Homo and Pan. Indeed, a parsimony analysis of the 53 segments and 37 protein coding genes leads to an estimate of Ne = 55,000 – 100,000. As this estimate is 5 to 10 times larger than the long-term effective population size of humans (~10,000) estimated from various genetic polymorphism data, the human lineage apparently had experienced a large reduction in effective population size after its separation from the chimpanzee lineage. Our analysis assumes a molecular clock, which is in fact supported by the sequence data used. Taking the orangutan speciation date as 12 to 16 million years (Myr) ago, we obtain an estimate of 4.8 to 6.4 Myr for the Homo-Pan divergence and an estimate of 6.5 to 8.7 Myr for the gorilla speciation date, suggesting that the gorilla lineage branched off 1.7 to 2.3 Myr earlier than the human-chimpanzee divergence.To study the genomic divergence between human and chimpanzee, large-scale genomic sequence alignments were performed. The genomic sequences of human and chimpanzee were first masked with the RepeatMasker and the repeats were excluded before alignments. The repeats were then reinserted into the alignments of non-repetitive segments and the entire sequences were aligned again. A total of 2.3 million base pairs (Mb) of genomic sequences, including repeats, were aligned and the average nucleotide divergence was estimated to be 1.22%. The Jukes-Cantor (JC) distances (nucleotide divergences) in non-repetitive (1.44 Mb) and repetitive sequences (0.86 Mb) are 1.14% and 1.34%, respectively, suggesting a slightly higher average rate in repetitive sequences.Annotated coding and noncoding regions of homologous and chimpanzee genes were also retrieved from GenBank and compared. The average synonymous and nonsynonymous divergences in 88 coding genes are 1.48% and 0.55%, respectively. The JC distances in intron, 5’ flanking, 3’ flanking, promoter, and pseudogene regions are 1.47%, 1.41%, 1.68%, 0.75% and 1.39%, respectively. It is not clear why the genetic distances in most of these regions are somewhat higher than those in genomic sequences. One possible explanation is that some of the genes may be located in regions with higher mutation rates.The major contributions of this study can be summarized as follows:(1) It demonstrates an effective and efficient method of fathoming genomic divergences between species without having to compare entire genomes.(2) It offers an important reference that can be applied to future studies of hominoid evolution.

Metrics

1 Record Views

Details

Logo image