Logo image
建立T3剖析樹語料庫:台語部分
Thesis

建立T3剖析樹語料庫:台語部分

劉亦真
Masters, 國立清華大學, 統計學研究所
2004

Abstract

剖析樹 treebank
T3 corpus is a treebank corpus consists of parallel sentences in the three major languages in Taiwan: Taiwanese, Hakka, and Mandarin. Those sentences are originally example sentences or phrases from “現代漢語八百詞”, and are translated into Taiwanese and Hakka by native speakers. The translated sentences are then segmented, Part-of-Speech tagged, and then syntactically bracketed; all done manually, although software tools are designed to help the laboring task of editing. This thesis reports the progress in the Taiwanese part. For the Part-of-Speech tagging, only few and somewhat incomplete literatures exist and we adopt a tag set of 23 tags. For bracketing, it seems that T3 corpus is the first treebank in Taiwanese. “Dotted tag” is a new system of tagging/bracketing Taiwanese phrase, and has the advantages of being easier to edit and explicitly promoting the use of phrase categories.

Metrics

1 Record Views

Details

Logo image