Logo image
Learning to Extract Aliases of Named Entities on the Web
Thesis

Learning to Extract Aliases of Named Entities on the Web

Hsieh, Hung-Ting
Masters, 國立清華大學, 資訊系統與應用研究所
2012

Abstract

關係抽取 別名辭典 專有名詞 網路語料庫 條件隨機域 Relation Extraction Alias Lexicon Named Entity Web as Corpus Conditional Random Field
A named entity (NE) can be referred to using many aliases in documents. Due to the prevalence of aliases, recognizing aliases of NE becomes an essential part for many applications such as Search Engine (e.g., Google, Bing) and Intelligent Dialog System (e.g., Siri). In this paper, we propose an approach for learning to extract aliases for a given NE on the Web automatically. The method involves generating applicable lexical patterns automatically so as to bias the search engine to return relevance documents containing aliases. Furthermore, we treat the process of identifying boundaries of aliases for a given NE as a sequence labeling problem and train a machine-learning model. At run-time, we bias the search engine to retrieve relevance snippets by transforming the given NE into a set of queries and then identify the boundaries of aliases with the trained model. We present a prototype, AliasFinder, which applies the method to find aliases from the Web. Experimental results show that the proposed method yields better performance than the baselines, provides an efficient way to find aliases of NEs.

Metrics

1 Record Views

Details

Logo image