Logo image
馬可夫語言模型應用di台語變調gah注音
Thesis

馬可夫語言模型應用di台語變調gah注音

洪俊詠
Masters, 國立清華大學, 統計學研究所
2004

Abstract

台語變調 注音 馬可夫語言模型 維特比搜尋 跨詞界變調 口語調
Taiwanese is rich in tone sandhi. It is a two-part problem: when to “tone sandhi,” and how to “tone sandhi.” For multi-syllabic words, major rules exist for both parts of the problems. For a complete sentence or a phrase consisting of multiple words, the tone sandhi rules for word may not apply at the last syllable of each word. Traditional approach to this problem is by the syntactic analysis, and this paper studies the tone sandhi problem by statistical approach. Using as corpora the seven volumes of Buddhist Sutra, published and phonetically annotated in Taiwanese by a senior nun, we model the phonetic transcription by syllable-based Markov language model, and study specifically the tone sandhi problem. A unigram model gives 80% correct and bigram model 84%. Both results are computed using seven-fold cross-validation.

Metrics

1 Record Views

Details

Logo image