Logo image
應用簡群最佳化演算法求解資料分群問題
Dissertation

應用簡群最佳化演算法求解資料分群問題

Lai, Chyh Ming
Doctor of Philosophy (PHD), 國立清華大學, 工業工程與工程管理學系
2015

Abstract

資料分群 簡群演算法 主成分分析 隨機取樣 K平均分群 K調和平均分群 data clustering K-means K-harmonic-means simplified swarm optimization principle component analysis random sampling technique
Data clustering is commonly employed in many disciplines. The aim of clustering is to partition a set of data into clusters, in which objects within the same cluster are similar and dissimilar to other objects that belong to different clusters. K-means (KM) and K-harmonic-means (KHM) are two common and fundamental clustering methods because of their simplicity and efficiency. However, both of them suffer from some problems. This study presents two novel algorithms based on simplified swarm optimization to deal with the drawbacks of KM and KHM, respectively. In addition, with the advance of internet and information technologies, the data size is increasing explosively and many existing clustering approaches including KM and KHM are inefficiency for dealing with the large-size problem. For that, we propose a clustering framework by exploring the connection between principle component analysis and a novel random sampling technique into a procedure to increase the scalability of the proposed clustering algorithm. To empirically evaluate the performance of the proposed methods, all experiments are examined using real-world datasets, and corresponding results are compared with recent works in the literature.

Metrics

1 Record Views

Details

Logo image