Abstract
This paper presents a case of lexical tone recognition for Mandarin speech using a combination of vector quantization and Hidden Markov Modeling techniques. The observation sequence was a sequence of vectorized parameters consisting of a logarithmic pitch interval and its first derivative. The vector quantization was applied to convert the observation sequence into a symbol sequence for Hidden Markov Modeling. The speech database was provided by 7 male and 7 female college students, with each pronouncing 72 isolated monosyllabic utterances. A probabilistic model for each of the four tones was generated. A series of tonal recognition tests were then conducted to evaluate the effects of pitch reference base, codebook size, and tonal model topology. Future consideration of Mandarin speech recognition is also discussed. © 1988 IEEE