Logo image
詞彙自動抽取及最佳化技術之研究
Thesis

詞彙自動抽取及最佳化技術之研究

張景新
Masters, National Tsing Hua University
1996

Abstract

精確率召回率最佳化詞彙抽取 precisionrecalloptimizationlexicon acquisition
本論文主要探討的問題是如何由大量文字中自動抽取新詞(new words;unknown words)或複合詞 (compound words);主要重點則在於如何有系統的改善精確率 (precision) 及召回率 (recall) 的綜合績效。為了能針對各不同語言在不同狀況下作最佳化處理,本論文分別就英文複合詞抽取及中文未知詞 (或新詞) 之抽取提出兩種不同層次的最佳化處理技術。一般英文複合詞抽取系統, 是根據候選詞的某些結合性特徵(associationfeatures),使用一個過濾器 (filter) 或分類器(classifier) 判斷該候選詞為 "詞" 或"非詞"。對於此類系統,本論文採用一套兩階段的最佳化策略,以使精確率及召回率的綜合績效達到最大或相對最大。第一階段的主要目的在於減少分類的錯誤率,第二階段的任務則在於,由第一階段所達成的最小錯誤率狀態出發,提昇精確率及召回率的綜合績效。為達到最小錯誤率的目標,本論文採用一系列的改善措施,有系統地改進錯誤率。同時,為達成第二階段的任務,本論文特別發展了一套可以依據使用者所指定的精確率及召回率函數,達到最大綜合績效的非線性學習策略。在中文未知詞抽取方面,一般系統會依據前後文的資訊 (contextual information), 作中文斷詞,再將候選詞送給分類器作判斷。由於有前後文訊息可資利用,因此,除了改變分類器的參數估計值,過濾掉一些不正確的字元組,來改善精確率之外,本論文也提出一套反覆式的方法,逐步找出相對於這些錯誤候選詞的正確新詞,以同時改善召回率;而不致因提高精確率而降低召回率,或反其道而行。 在抽取英文複合詞方面,本論文所提的技術在抽取二元 (bigram) 複合詞時,精確率及召回率的加權平均值 (weighted precision-recall, WPR) 可達88% 左右,三元(trigram) 複合詞的 WPR 亦達 88% 左右。另一綜合績效指標 F-measure,則分別為 84%(bigram) 及 86% (trigram) 左右。(語料由 20,715 句訓練語料及 2,301 句測試語料所組成)。 採用本論文所提的最佳化學習技術時,亦可明顯觀察到精確率及召回率隨指定的績效指標而變動的情形。 在中文未知詞抽取方面,精確率及召回率大致呈單調遞增的趨勢。而不像一般系統,在精確率提高時召回率會降低,或發生相反情況的現象。在測試311,591 句中文語料時,所得到的 F-measure分別為 76% (bigram), 54% (trigram) 及70% (quadgram). 與採用非反覆式方法所獲的的績效 74% (bigram), 46% (trigram) 及58%(quadgram) 相比,有相當顯著的改善。Automatic lexicon acquisition from large text corpora issurveyed in thisdissertation, with special emphases onoptimization techniques for maximizingthe joint precision-recall performance. Both English compound word extractionandChinese unknown word identification tasks are studied in orderto exploreprecision-recall optimization techniques indifferent languages of differentcomplexity using differentavailable resources. In the English compound wordextractiontask, the simplest system architecture, which assumes thatthelexicon extraction task is conducted using a classifier (or afilter) based ona set of multiple association features, isstudied. Under such circumstances,a two stage optimizationscheme is proposed, in which the first stage aims atminimizingclassification error and the second stage focuses onmaximizingjoint precision-recall, starting from the minimumerror status. To achieveminimum error rate, variousapproaches are used to improve the error rateperformance of theclassifier. In addition, a non-linear learning algorithmisdeveloped for achieving maximum precision-recall performancein terms of userspecified objective function of precision andrecall. In the Chinese unknownword extraction task, wherecontextual information as well as word associationmetrics areused, an iterative approach, which allows us to improvebothprecision and recall simultaneously, is proposed toiteratively improve theprecision and recall performance. Forthe English compound word extractiontask, the weightedprecision and recall (WPR) using the proposed approachcanachieve as high as about 88% for bigram compounds,and 88% for trigramcompounds for a training (testing) corpusof 20715 (2301) sentences sampledfrom technical manuals of cars.The F-measure performances are about 84% forbigrams and 86% fortrigrams. By applying the proposed optimization method,theprecision and recall profile is observed to follow the preferredcriteriaof different lexicographers. For the Chinese unknownword identification task,experiment results show that bothprecision and recall rates are improvedalmost monotonically,in contrast to non-iterative segmentation-merging-filtering-and-disambiguation approaches, which often sacrifice precisionforrecall or vice versa. With a corpus of 311,591 sentences,the performance is76% (bigram), 54% (trigram), and 70%(quadgram) in F-measure, which issignificantly better thanusing the non-iterative approach with F-measures of74%(bigram), 46% (trigram), and 58% (quadgram).

Metrics

1 Record Views

Details

Logo image