Logo image
語音辨識 輔助的 台語語料庫 收集方法 探討
Thesis

語音辨識 輔助的 台語語料庫 收集方法 探討

游聲峰
Masters, National Tsing Hua University
2013

Abstract

Corpus collectionSpeech recognition
Corpus is fundamental to computing linguistics. But for marginalized Taiwanese language, corpus collection is not as easy as Chinese. This thesis explores using speech recognition technology to help collect Taiwanese text and speech corpus with various annotations.Given a Taiwanese sentence and its corresponding recorded speech, we might semi-automatically obtain its phonetic annotations and tone sandhi. This gives a total of four corpus contents: text, speech, phonetic annotation, and tone sandhi. Let us call it Taiwanese-text-Taiwanese-speech (TTTS) problem. Another similar setup is the Mandarin-text-Taiwanese-speech (MTTS) problem. In addition to the four corpus contents, we might also obtain Taiwanese Mandarin parallel sentences in the MTTS case. Parallel corpus is essential to the research of Taiwanese-Mandarin translation.Since the current automatic speech recognition system is not perfect yet even for healthy languages like English and Chinese, it is sensible to manipulate the recognition network to decrease the complexity of the network used in the speech recognition system. Using a TTTS corpus and a MTTS corpus, this paper explores ways of constructing the recognition network on a sentential basis both for Taiwanese text and for Mandarin text. The current hidden Markov model based speech recognition system is capable of giving two kinds of results. One is the best path in the recognition network, in the likelihood sense. The other is the occupation time of each syllable. These results can be used in spottin possible errors in the corpus.

Metrics

1 Record Views

Details

Logo image