Logo image
利用Spline內插技術改進時序基因資料的AP分群方法研究
Thesis

利用Spline內插技術改進時序基因資料的AP分群方法研究

Hsu, Ting-Chieh
Masters, 國立清華大學, 資訊工程學系
2009

Abstract

Spline 內插技術 親和性互動式演算法 基因表現時間序列 Spline Interpolation Affinity Propagation Time-Series Gene Expression
DNA microarray technology has been widely used in life science research for many years. The technology allows scientists monitoring genes' expression level during biological processes simultaneously. Analyzing massive time-series data is important to explore the complex dynamics of biological systems. However, the analysis task of time-series gene expression data is difficult since noise levels and measurement uncertainties are high. The early clustering methods such as k-means, self-organizing maps and hierarchical clustering disregarded the temporal dependency between successive time points. As for probabilistic model-based methods, dynamic Bayesian networks (DBN) and hidden Markov models (HMM), are more suitable for time-series but fail in computational inefficiency. In addition, real gene datasets has undersampling problem for long intervals between time points of harvesting expression data. In this thesis, an unsupervised clustering algorithm which combines Spline interpolation and Affinity Propagation is proposed. The proposed method investigates the relationship between genes across distinct time points through the interval selection after using interpolation to eliminate the influence of undersampling. We demonstrate our method result in significant accuracy on real gene expression time-series datasets without \textit{priori} knowledge such as the number of clusters and exemplars. Our study provides a way of clustering gene expression time-series data for future biological investigations.

Metrics

1 Record Views

Details

Logo image