Abstract
將相似的事物匯聚,劃分成為一族群,俾使有所區分且容易辨識,這是人類經常從事的基本活動。而聚類分析法正是主要的一種科學方法來解決此領域的問題,其兩個目標分別在求得分類後的族群內具有同質特性與族群間適當的隔離。然而,在分類過程中,族群的界定常具有不確定的特性,所以,乏晰集合被引進,使得本類問題得以有效分析。聚類分析法涉及許多層面問題,分類方法則是其中的主要一環。至今,許多不同的分類方法已相繼提出,其中目標函數法最易於洞察問題的結構,乃為本研究採用的方法。唯現有分類方法的理論基礎多考量同質的特性,因此,本研究的第一項工作是要建立一個雙目標數學規劃的模式,以符合聚類分析法的需要。關於``同質''的目標函數,我們以距離與統計兩種觀點,評估文獻中既有的函數,獲得兩項函數引進本模式,分別為FCM 與FMLP,我們並改善FMLP在原文獻中求解過程的瑕疵。另外,關於隔離的目標函數,我們提出三種不同觀點的函數,經評估後只有一項適合導入,其為族群中心點間距離平方和。其它聚類分析法涉及到的爭議性問題,諸如特徵項的選擇、接近度的訂定、族群數量的裁決等等,在本研究也將一併探討並提出解決方案。總結理論上的研討,我們建議一套有系統的整合流程來完成聚類分析法,範圍包括從抽樣到鑑識分類,這是本研究的第二項工作。最後,本研究建議的方案要以中興大學王徵教授所提供的2899筆臺灣靈芝樣本作實際上的測試,以探討建議方案可行性。結果有12種臺灣靈芝分類被提出,與現有分類系統不吻合率為33%,但其中資料誤差占了29.1%,而模式誤差僅占 3.9%。由實際的測試中,驗證了雙目標分類模式確實符合本研究擬改善的目標,同時,我們建議的乏晰聚類分析系統流程也堪稱完備。本文分為五章,第一章介紹提要。第二章則作文獻回顧,聚類分析法涉及到的爭議性問題在此一一被討論。在第三章中,我們建立一個雙目標數學規劃的模式來達成聚類分析的目標,並且建議一套有系統的資料流程來完成聚類分析法。本研究建議的方案在第四章中,以臺灣靈芝樣本作實際上的測試。本研究的總結在最後一章提出,對臺灣靈芝分類的建議以及未來值得後續探討的部分也一併提出。Normally, cluster analysis is to find homogeneous and well-separated subsets. However, because vague boundaries of theclusters usually occur in practice, introducing fuzzy settheory has been suggested to handle this porblem. Clusteranalysis involves several research issues and clustering methodis the main issue. So far, many clustering methods have beendeveloped where the type of objective function methods canprovide an insight into the structure of the problem and thusis adopted in our study. Since the existing methods focus onthe homogeneous property and neglect the well-separated one, todevelop a model of which both criteria are taken intoconsideration is thus our first task in this study. Forhomogeneous criterion, models FCM and FMLP are reviewed andevaluated from both distance and statistic viewpoints and theFMLP approach is improved in this thesis. For well-separatedcriterion, three feasible functions are proposed for comparisonand the function of total distances between cluster centers isadopted. Other issues in cluster analysis with their solutionmethods are also discussed, such as the choice of attributes,the measurement of the closeness and determination of number ofclusters. After all, a systematic process of fuzzy clusteranalysis is presented. This task including sampling andclassification is our second aim of the study. Finally, theproposed method is tested and evaluated with the case of TaiwanGanoderma that is provided by Prof. Chen Wang with 2899 sampledata in total. On this case, we suggest twelve patterns ofTaiwan Ganoderma and, based on current taxonomy, the unmatchingrate is 33% which includes the data error rate of 29.1% and themodel error rate of 3.9%. It is concluded that the proposedclustering model can serve our purposes of the study and thesystematic process operates reasonably.