Logo image
T3台語剖析樹語料庫與Brill剖析器
Thesis

T3台語剖析樹語料庫與Brill剖析器

史滌明
Masters, 國立清華大學, 統計學研究所
2005

Abstract

T3台語剖析樹語料庫 Brill剖析器 台語詞組結構 基於轉換規則的錯誤驅動分析 K-Fold交叉驗證
T3 corpus is a treebank corpus consists of parallel sentences in the three major languages in Taiwan: Taiwanese, Hakka, and Mandarin. Those sentences are originally example sentences or phrases from “現代漢語八百詞”, that are translated into Taiwanese and Hakka by native speakers. The translated sentences or phrases are then segmented, part-of-speech tagged, syntactically bracketed, and furtherly annotated with structure type for all the immediate constituents. All works are done manually with help by the “T3bracket”, a Windows program specifically designed for this task. Despite of studying various materials before we embark this task, we are still faced with many difficulties in all the phases of translation, segmentation, and bracketting. And thus, discussions was help regularyly for a period of two years. The Taiwanese part of the T3 treebank is almost finished; double check is still required. The transformation-based error-driven parsing of Brill(1993) is applied to part of the T3 treebank. The resulting parser, has non-crossing bracketting accuracy 87.8% for inside test and 89% for outside test.

Metrics

1 Record Views

Details

Logo image