Abstract
In the past four decades, the integration of Hidden Markov Model into automatic speech recognition (ASR) has made great progress. A variety of research about HMM in speech recognition has been developed, such as different definitions of observation probability functions, or different methods for estimating model parameters. Recently, some researchers looked into the framework of ASR from a new angle. C.H. Lee proposed automatic speech attribute transcription (ASAT) which combines the knowledge-based and data driven methods. Some diagnostic information is provided by integrating additional speech attribute detectors which are designed by acoustic phonetic knowledge. It is believed that incorporation of such knowledge is potentially beneficial to ASR. ASAT is based on speech attribute detection. Specific speech event is composed of some speech attributes. ASAT can interpret different speech event into more high level speech evidence. Therefore, effective speech attribute detectors have become an important research issue. In this paper, we focus on detecting articulation attributes, namely, vowel, fricative, stop, sonorant consonant and silence. We use knowledge-based features to detect manner of articulation. Support vector machine based classifiers for manner of articulation have also been designed using a set of knowledge based features under a probabilistic framework. Besides frame based event detectors, segment based detectors can also be used. Some speech landmarks are detected by using temporal information. These landmarks are integrated with HMM-based event detectors.