Abstract
The long-term goal of this research is to construct a Mandarin-English bilingual speech recognition system on devices mounted on automobiles with limited storage size. Thus, the purpose of this thesis is to effectively reduce the model size and to maintain considerable performance as a unilingual system without using language identification. In this thesis, similar acoustic models are merged to reduce the number of model parameters. Similar acoustic units between the two languages are found by analyzing different phonetic notations with either knowledge-driven or data-driven techniques. In addition to directly merging the two acoustic models, this thesis also proposes the use of decision trees to merge states of different HMMs (hidden Markov models). Experimental result shows that, merging the models in a finer level via decision trees not only effectively reduces the model size but also enhances robustness of the bilingual models. By comparing to the baseline models, the state mergence using decision trees can reduce model size to one third of the original one and achieve an improvement of 1.2% in correction rate of bilingual recognition.