Abstract
This thesis examines how pronunciations of disyllabic words vary across word-internal syllable boundary for spoken Mandarin Chinese and how multiple pronunciation dictionaries constructed via the proposed variant-selection algorithm may significantly enhance the performance of pronunciation modeling for speech recognition. Three preprocessing stages prior to the selection of typical variants in multiple pronunciation dictionaries are: 1) to derive word variants by a free phone recognizer; 2) to calculate similarity scores by aligning phonetic surface forms to canonical forms; 3) to categorize pronunciation variants by stipulated reduction rules on the changes of the word-internal syllable structure. In our work, four reduction types are suggested by considering the presence of a within-word syllable boundary: Citation form-like reduction, marginal segment deletion, nuclei merger, and syllable merger. Additionally, the results on a series of statistical analyses show that the most frequent reduction types for disyllabic words in Chinese conversation are citation form-like reduction and syllable merger. In particular, high-frequency disyllabic words preferentially take the extreme syllable-merger form. Furthermore, our results show that segmental reduction in Chinese disyllabic words is morphology-dependent, and also related to the prosodic position at which a disyllabic word is produced as well as the temporal quality of the word. Motivated by the quasi-categorical reduced forms of disyllabic words produced in Chinese conversational speech, a frequency-based selection procedure is adopted to select typical pronunciations in the dictionary. The implementation of our new pronunciation models derived from the training data has shown a 2.4% absolute improvement on the domain-specific recognition task (MMTC), and an enhancement by 1.2% on the freely-conversed recognition task (MCDC8). Even though the confusability in provided lexicons is increased, our findings suggest that the automatically learned pronunciation models may capture more linguistic variation beyond short-span contextual effects, such as phoneme substitution and deletion.