Logo image
可調式之中文文件自動摘要
Thesis

可調式之中文文件自動摘要

陳鈺瑾
Masters, National Tsing Hua University
1999

Abstract

可調中文摘要 scalablesummarizationChinese text
This paper proposes an approach to generate scalable summaries for Chinese text automatically. We observe that summaries usually consist of topic sentences, and topic sentences usually contain topic phrases. Chinese words are not like English ones, which are separated by white spaces, therefore we have to carry out word segmentation before identifying topic phrases.We adopt a dynamic programming method based on Markov Model to segment and tag words for known as well as unknown words. Then, we identify topic phrases of the article based on linguistic properties of topic phrases at syntactic and discourse levels. At syntactic level, the topic phrases always follow a limited set of syntactic patterns, while at discourse level, the topic phrases always repeat in the article. After identifying topic phrases, we divide the article into subtopic segments, because authors often divide one subtopic into several paragraphs for readability. We merge the most similar adjacent sentences into one segment using clustering method, and extract one sentence from each segment as summary. We design six scoring methods to calculate the imformativeness of sentences, including measurement with topic phrase length, topic phrase frequency, and topic phrase count, with or without lead weight. The experiment illustrates that the lead weight methods perform better, and among all scoring methods, the measurement with topic phrase length has the best performance. In the future, we will combine more article features such as cue phrase to produce better results. Further more, we can shorten or combine the sentences using some reduction and combination rules to produce summaries with quality approaching the manual ones.

Metrics

1 Record Views

Details

Logo image