Logo image
A. Prediction by zone and its application to egg productivity in chickens B. A simple genotype imputation method and copy number haplotype inference with Hidden Markov Model and localized haplotype clustering
Dissertation

A. Prediction by zone and its application to egg productivity in chickens B. A simple genotype imputation method and copy number haplotype inference with Hidden Markov Model and localized haplotype clustering

林沿妊
Doctor of Philosophy (PHD), 國立清華大學, 資訊工程學系
2011

Abstract

區域預測法 雞蛋的生產力 基因型的補差法 單體型的複製數 隱馬可夫模型 單體型 Prediction by zone egg productivity chickens genotype imputation copy number haplotype Hidden Markov Model haplotype
A. Prediction by zone and its application to egg productivity in chickens Taiwan red-feathered country chickens (TRFCCs) are one of the main meat resources in Taiwan. Due to the lack of any systematic breeding programs to improve egg productivity, the egg production rate of this breed has gradually decreased. The prediction by zone (PreZone) program was developed to select the chickens with low egg productivity so as to improve the egg productivity of TRFCCs before they reach maturity. Three groups (A, B and C) of chickens were used in this study. Two approaches were used to identify chickens with low egg productivity. The first approach used predictions based on a single dataset, and the second approach used predictions based on the union of two datasets. The levels of four serum proteins, including apolipoprotein A-I, vitellogenin, X protein (an IGF-I-like protein) and apo VLDL-II, were measured in chickens that were 8, 14, 22 or 24 weeks old. Total egg numbers were recorded for each individual bird during the egg production period. PreZone analysis was performed using the four serum protein levels as selection parameters, and the results were compared to those obtained using a first-order multiple linear regression method with the same parameters. The PreZone program provides another prediction method that can be used to validate datasets with a low correlation between response and predictors. It can be used to find low and improve egg productivity in TRFCCs by selecting the best chickens before they reach maturity. B. A simple genotype imputation method and copy number haplotype inference with Hidden Markov Model and localized haplotype clustering High-throughput technology for genotyping has made genome-wide associations possible. Single nucleotide polymorphism (SNP) data derived from array-based technology are usually flawed due to missing data, although they have generally high call rates and good concordance rates across different genotype calling schemes. Missing SNPs can bias the results of association analyses and hence loci with missing data are removed in some studies. Imputation is a method of compensating for the missing data by filling in the most probable values. It can increase the power of the association study and does not involve extra cost to genotype the missing SNPs. In this article, we propose a simple imputation method (Simpute) that takes advantage of the high resolution of SNPs in either the array platform or the mass parallel sequencing platform. It is based on the linkage disequilibrium (LD) structure of the chromosome and only two nearby SNPs are needed to fill in the missing data. Simpute does not use any reference data. Simpute provides a simple, accurate and fast solution to the whole genome imputation. We have demonstrated that when the SNPs are densely distributed on the chromosome with high linkage disequilibrium between adjacent loci, there is no need to adopt complicated algorithms. Simpute is suitable for regular screening of the large scale SNP genotyping especially when the sample size is large and the efficiency is a major issue of the workflow. Copy number polymorphisms and aberrations can now be studied at high resolution using genome-wide SNP arrays. This paper presents a method based on Hidden Markov Model to detect parent specific copy number change on both chromosomes. A haplotype tree is constructed with dynamic branch merging to model the transition of the copy number status of the two alleles assessed at each SNP locus. The emission models are constructed for the genotypes formed with the two haplotypes. The proposed method can provide the segmentation points of the copy number variation regions as well as the haplotype phasing for the allelic status on each chromosome. The estimated copy numbers are provided as fractional numbers, which can effectively accommodate the somatic mutation in cancer specimens that usually consist of both normal and mutant cells. The algorithm is evaluated on the previously published regions of copy number variation on the 270 HapMap individuals. The results were compared with five popular methods: PennCNV, genoCN, COKGEN, QuantiSNP and cnvHap. The proposed algorithm exhibits roughly comparable sensitivity of the CNV regions to the best algorithm in our genome-wide study and demonstrates the highest detection rate in SNP dense regions. In addition, we provide better haplotype phasing accuracy than similar approaches.

Metrics

1 Record Views

Details

Logo image