Logo image
Information Extraction and Classification for Person Search
Thesis

Information Extraction and Classification for Person Search

Jie-Fei Yang
Masters, 國立清華大學, 資訊系統與應用研究所
2005

Abstract

人名檢索 資訊擷取 文件分類 person search information extraction text categorization
We introduce a method for automatically collecting personal information and professional domain of the person. In our approach, personal information is extracted and the domain is identified from web-based data based on personal name disambiguation. In the training phase, the method involves generating surface pattern to personal information extraction based on linguistic and statistical information from the Web, and an unsupervising algorithm for constructing Web-based text categorization. At runtime, submitting a person name into a search engine, extracting personal information and identifying each retrieved passage the domain according to the expected person name, finally the referents are sorted by domain, personal information and the degree of popularity. We also described an implementation of the proposed method. Blind evaluation of a set of names shows that our method outperforms extracting personal information and cleanly classifying individual’s domain-specific knowledge. This method can be applied to help users quickly find about a person with resulting in the display of personal information in a systematic and consistent way.

Metrics

1 Record Views

Details

Logo image