Logo image
延複詞及延複詞類初探
Thesis

延複詞及延複詞類初探

陳建忠
Masters, 國立清華大學, 統計學研究所
2009

Abstract

延複詞 延複詞組 斷詞 分詞 讓格書寫 extended compound word extended compound part-of-speech segmentation LangGeh orthography
With traditional orthography in Chinese or Taiwanese where the writing is without spaces between words, segmentation is both fundamental and difficult. It is difficult because there hardly are clear boundaries between words and compound words, and between words and phrases. The current segmentation standard proposed in CKIP (1996) relies mainly on semantics and syntaxes, and noticeably gives inconsistent segmentation results. On the other hand, we find that the literal forms of character strings are much easier to recognize. We thus propose segmentation in so-called extended words. We emphasize the use of six literal forms to define extended words: 1. character repetition patterns, 2. two-character strings are loosely defined as an extended words, 3. noun in 2+1 shape with head word at the right, 4. concatenation of parallel words, 5. total number of words, 6. total number of constituents. Extended words include simple extended words and general extended words. Simple extended words correspond roughly to units segmented by CKIP standard, while using much simpler rules. The general extended word consists of multiple constituents with total length up to four or five characters, while keeping syntactic structure simple. Interestingly the segmentation in extended words and the LangGeh orthography (江永進等(2009)) give similar results. We also try to tag the extended words with syntactic categories. Due to the fact that we use a larger unit, we are given the opportunity to omit tagging those constituents of single character which are syntactically complex, and results in a simpler tagging process.

Metrics

1 Record Views

Details

Logo image