Abstract
In this thesis, we propose two new techniques of reducing and estimating error rate for robust speech recognition. A robustness technique, SNR-incremental stochastic matching (SISM) algorithm, is proposed to reduce the mismatch between the training and testing conditions by some form of compensation and consequently to improve the recognition performance. The goal of error rate estimation is for monitoring the recognition performance such that the appropriate schemes for improving the recognition accuracy can be performed according to the performance degradation. The SISM algorithm is an extension of Sankar and Lee’s stochastic matching (SM) for dealing with the distortion due to additive noise. We address two issues concerning the original maximum likelihood-based SM techniques. One concern is that the initial condition of the expectation-maximization (EM) algorithm has to be set carefully if the mismatch between training and testing is large. The other is that the performance is often limited by the newly adapted model in noise compensation instead of reaching the higher level of accuracy often obtained in clean environments. Our proposed SISM algorithm attempts to improve the initial condition and to relax the performance bound. First, the SISM algorithm provides a good initial condition making use of a set of environment-matched models. The second is a recursive operation, i.e. the reference model in each recursion is changed along the direction of SNR increment in order to push the recognition performance to that obtained at higher SNR levels. Experimental results show that the SISM algorithm provides further improvement after the best environment-matched performance has been reached, and can therefore obtain an additional discriminative power through using the speech models with higher SNR instead of retraining process. A model-based error rate estimation framework is proposed for speech and speaker recognition. It aims at predicting the performance of a hidden Markov model (HMM) based recognition system for a given task vocabulary and grammar without the need of running recognition experiments using a separate set of testing samples. This is highly desirable both in theory and in practice. However, the error rate expression in HMM-based speech recognition systems has no closed form solution, due to the complexity of the multi-class comparison process and the need for dynamic time warping to handle speech patterns of different sizes and of various lengths. To alleviate the difficulty, we propose a one-dimensional model-based misclassification measure to evaluate the distance between a particular model of interest and a combination of many of its competing models. The error rate for a class characterized by the HMM is then the value of a smooth zero-one error function given the misclassification measure. The overall error rate of the task vocabulary could then be computed as a function of all the available class error rates. The key here is to evaluate accurately the misclassification measure without using any speech data. In this paper, we show how the misclassification measure could be approximated by first computing the distance between two mixture Gaussian densities, then between two HMM’s with mixture Gaussian state observation densities and finally between two sequences of HMM’s. When comparing the error rate obtained in running actual experiments and that of the new framework without using any test data, the proposed algorithm accurately estimates the error rate of many types of speech and speaker recognition problems. Based on the same framework, it is also demonstrated that the error rate of a recognition system in a noisy environment could also be predicted.