Logo image
Linguistically Motivated Word Pronunciation Modeling for Automatic Speech Recognition of Chinese Conversational Speech
Dissertation

Linguistically Motivated Word Pronunciation Modeling for Automatic Speech Recognition of Chinese Conversational Speech

Liu, Yi-Fen
Doctor of Philosophy (PHD), 國立清華大學, 資訊系統與應用研究所
2015

Abstract

語音變體 音節縮讀 語音弱化類型 字詞類型 發音變異模型 雙音節字詞 自然口語語音辨識系統 Pronunciation variation Syllable contraction Reduction type Word type Pronunciation modeling Disyllabic word Spontaneous Speech Recognition ASR system
This thesis examines how pronunciations of disyllabic words vary across word-internal syllable boundary for spoken Mandarin Chinese and how multiple pronunciation dictionaries constructed via the proposed variant-selection algorithm may significantly enhance the performance of pronunciation modeling for speech recognition. Three preprocessing stages prior to the selection of typical variants in multiple pronunciation dictionaries are: 1) to derive word variants by a free phone recognizer; 2) to calculate similarity scores by aligning phonetic surface forms to canonical forms; 3) to categorize pronunciation variants by stipulated reduction rules on the changes of the word-internal syllable structure. In our work, four reduction types are suggested by considering the presence of a within-word syllable boundary: Citation form-like reduction, marginal segment deletion, nuclei merger, and syllable merger. Additionally, the results on a series of statistical analyses show that the most frequent reduction types for disyllabic words in Chinese conversation are citation form-like reduction and syllable merger. In particular, high-frequency disyllabic words preferentially take the extreme syllable-merger form. Furthermore, our results show that segmental reduction in Chinese disyllabic words is morphology-dependent, and also related to the prosodic position at which a disyllabic word is produced as well as the temporal quality of the word. Motivated by the quasi-categorical reduced forms of disyllabic words produced in Chinese conversational speech, a frequency-based selection procedure is adopted to select typical pronunciations in the dictionary. The implementation of our new pronunciation models derived from the training data has shown a 2.4% absolute improvement on the domain-specific recognition task (MMTC), and an enhancement by 1.2% on the freely-conversed recognition task (MCDC8). Even though the confusability in provided lexicons is increased, our findings suggest that the automatically learned pronunciation models may capture more linguistic variation beyond short-span contextual effects, such as phoneme substitution and deletion.

Metrics

1 Record Views

Details

Logo image