Logo image
半督導式中文特定類型具名實體擷取之研究
Thesis

半督導式中文特定類型具名實體擷取之研究

陳瑩綺
Masters, National Tsing Hua University
2008

Abstract

資料擷取具名實體辨識網路語料庫最大熵模型自動標記 Information extraction (IE)Name Entity Recognition (NER)Web corpusMaximum Entropy model (ME)Automatically tagging
We introduce a semi-supervised method for the extraction of instances of a certain type from a Chinese text under a domain. In our approach, a machine learning model for extraction is trained on an automatically collected and tagged corpus, aiming at eliminating the limiting factor of human annotation on current supervised systems. The method involves selecting seed data of target instances from off-the-shelf general purpose thesauri, using seeds to automatically collect a corpus from the Web, automatically tagging the corpus by seed data and training a machine learning model on the corpus. At run time, a natural language text is segmented into words, and the trained model is applied on the words to make the best tagging decisions, from which we extract target instances. The evaluation of exact match on a set of annotated test data shows that the method successfully extracts target instances at the precision rate of 78%. Our methodology accomplishes the elimination of human annotation on training data by small amount of seed data, and the method is highly portable to other domains.

Metrics

1 Record Views

Details

Logo image