Logo image
T3台語剖析樹語料庫與Brill詞類標記
Thesis

T3台語剖析樹語料庫與Brill詞類標記

周思源
Masters, 國立清華大學, 統計學研究所
2005

Abstract

詞類標記 N-gram語言模型 馬可夫模型 隱藏馬可夫模型 維特比演算法 K-Fold交叉驗證 Brill詞類標記 Part-of-Speech Tagging N-gram Markov language Model Hidden Markov Model Viterbi Algorithm Deleted Interpolation K-fold Cross Validation Transformation-Based Error-Driven Learning
Part-of-Speech Tagging is a basic issue in the natural language processing. In this paper, we study the effect of Brill Tagger (1992) using part of the T3 Taiwanese treebank. Brill tagger is a transformation-based error-driven approach. Based on the results of other tagging method such as N-gram language model, Brill tagger learns a set of transformation rules from an annotated corpus. The learning process is error-driven in that its objective is to minimize the tagging errors computed from the comparison of the transformed results to the standard annotated corpus. Annotated corpus is often suffered from inconsistency problem, and we also study the problem using the confusing matrix. The best tagging result that we obtained is 92.80% and 85.59% for the inside test and the outside test respectively.

Metrics

1 Record Views

Details

Logo image