Logo image
詞彙語意歧義解析-適應性的概念式作法
Thesis

詞彙語意歧義解析-適應性的概念式作法

陳振南
Masters, National Tsing Hua University
1997

Abstract

語意歧義解析自然語言處理機讀字典字義分類適應性做法資訊檢索 word sense disambiguationnatural language processingmachne-readable dictionarysense divisionadaptive approachinformation retrieval
中文摘要字意歧義的問題是各種不同範疇之自然語言處理應用中,最基本的課題。研究者通常運用大型語料建立詞彙語意知識庫,這種作法所建立的語言模式通常只能處理特定題材之文章。由於同一個詞彙在不同性質之語料會有不同之字意解讀,因此如何獲得均衡性及完整性之語意知識是語料庫作法所面臨之考驗與挑戰。一般而言,語意知識的建立過程中,完整性的問題一直是知識蒐集的瓶頸,而字典可視為一個中性語料庫因它具有完整性之語意知識且同一個詞彙之語意區隔相當細緻。本論文主要從字典建立一自動化的語意知識模型,首先從字典的定義及例句所使用之詞彙建立基本語意知識,爾後分析字典內的抽象語意知識並標示相關之主題以擴充基本語意知識。另外我們提出一個辨識測試文章內文脈難易度的分析模型,藉由所建立之語意知識先處理文脈中簡單的歧義問題,接著再蒐集這些蘊藏在簡單型文脈內的語意知識用以調適先前用字典所建構之基本語意知識,這種由簡單型文脈所建立之語意知識可有效處理文章中其他較難的字意歧義問題。實驗顯示,從字典可有效提煉出相當於大型語料所蘊藏之語意知識,而這些語意知識可依不同題材之測試語料加以自動調適,這種調適性之作法開拓了自然語言處理中解決字意歧義的另一研究空間。AbstractWord sense disambiguation for unrestricted text is oneof the most difficult tasks in the fields of computationallinguistics. The crux of the problem is to discover a modelthat relates the intended sense of a word with its context.Such relations allow for other computational semantics taskssuch as noun sequence interpretation and prepositional phraseattachment disambiguation. There are two aspects of building upa knowledge base of semantic relations. First, we need arepresentational scheme to divide and codify the word senses ina context. Both word-based and class-based representation ofword senses and context have been used in the literature.Second, based upon the sense division, an algorithm is used toidentify semantic relations in a lexical resource, a corpus or amachine-readable dictionary. Until recently, most knowledgebases are created manually by language experts. The awakeningof the statistical approach did provide an alternative ofacquiring a knowledge base from corpora. In recent years,attention has been shifted from word-based towards class-basedrepresentation, in the hope of providing broader coverage forunrestricted text. The goal of this dissertation is to makesignificant progress on constructing a class-based sense modelfrom existing resources that is effective in representing thesemantic relations and resolving sense ambiguity. The modelshould have a broad coverage of senses making efficientrepresentation of semantic relations possible. In particular,we describe a series of algorithms based on informationalretrieval techniques that cluster machine-readable dictionary(MRD) senses to provide a complete and appropriate sensedivision for word sense disambiguation (WSD). One algorithmexploits the topical sense clusters available in Longman'sLexicon of Contemporary English. Another algorithm identifiesthe topics related to terms in the definition of a dictionaryheadword. In other words, senses are classified according toeither their topics or disambiguated genus. Therefore, our maintool in building up a sense division, as well as semanticrelation, is topical analysis of dictionary definitions. Wealso describe a general framework for an adaptive conceptualword sense disambiguation. The learning process described herebegins with an initial disambiguation step of knowledge based onMRDs. An adaptation step follows to combine the initialknowledge base with knowledge gleaned from the partiallydisambiguated text. Once the knowledge base is adjusted to suitthe text at hand, it is then applied to the text again tofinalize the disambiguation result. Definitions and examplesentences from Longman Dictionary of Contemporary English(LDOCE) are employed as training materials for WSD, whilepassages from the Brown corpus and Wall Street Journal are usedfor testing. Finally, we report on several experimentsillustrating the effectiveness of the adaptive approach.

Metrics

1 Record Views

Details

Logo image