Abstract
Sentence-end vowel devoicing and bidakuon allophones are common problems in Japanese speech recognition. This thesis proposes the use of specialized models for sentence-end vowel phones to overcome the devoicing problem and an automatic transcription correction framework for bidakuon allophones. In this study, Mel-frequency cepstral coefficients (MFCC) and log energy are used as features for training speech recognition models. Sentence-end vowel models are adopted for each sentence during the training phase in order to improve the recognition performance at the end of the sentence. On the other hand, we use an automatic transcription correction framework to resolve the bidakuon allophone problem by an iterative correction method. The iterative correction method is based on thresholds trained from the ranking scores. The transcription is corrected gradually towards the actual pronunciation recorded in the training data. We use three types of performance measure to evaluate the effectiveness of the proposed methods. They are confidence measure based on phone model ranking, free-mola decoding, and sentence recognition. The experimental results show that using both of the proposed methods can effectively enhance the recognition performance of the baseline system.