Logo image
基於背景資訊與篇章結構來解析小說故事中的代名詞
Thesis

基於背景資訊與篇章結構來解析小說故事中的代名詞

Ritika Nimje
Masters, 國立清華大學, 資訊工程學系
2016

Abstract

背景資訊 故事 代名詞 pronoun resolution fictional stories background information
Pronoun resolution is a well-known task in discourse analysis and is an important research issue in the applications of natural language processing. We tackle the problem of pronoun resolution in the textual content by leveraging background semantic information of the characters in the story and the discourse structure. We also extracted some general discourse rules from the story about the speaker of the dialogs to split the story text into clusters. Background information includes, the relationship between the main characters in the story. We extracted some general discourse rules from the story about the narrator and the speakers of dialogs to split the story text into clusters. Specifically, we focus on noun phrases that co-reference identifiable entities that appear in the text; the challenge in this context is to improve the pronoun co-reference resolution by leveraging the potential relations by which we can identify the mentions. Our system applies state-of-the-art techniques to extract entities, noun phrases, and candidate co-references that are conducted by the Stanford Parser’s co-reference resolution method. Since Stanford parser’s co-reference resolution can account for only about 21% to 35%~20.9%, ~-27.985% and ~34.98% accuracy of pronoun resolution we need, we propose an augmented approach in which we assume we could provide priorly and self (manually) annotated data (which is about 10% of a full text) to Stanford parser and utilize the semantic relatedness of noun phrases to the background information about the characters (it included the person relationship like “father”, “mother”, “daughter”, etc. about the characters) to resolve the co-references. We employ heuristic rules of splitting text into segments based on the discourse structure as well as the background information to improve the recall and precision of pronoun resolution in stories. We use three stories with different domains and different writing styles, we used “A Great Gatsby”, “A Case of Identity” and “Harry potter” in our experiments and got, there was about ~50% of improvement in precision after applying ourt methods.

Metrics

1 Record Views

Details

Logo image