Abstract
This research focuses on selecting a specific training data from the existing corpus to conduct a discriminative feature transformation with existing hidden Markov models (HMM) to improve the performance of task-specific English speech recognition. This thesis contains two parts: the first part is task-specific corpus selection for discriminative feature transform using heteroscedastic linear discriminant analysis (HLDA); the second part is feature mergence. Two methods are used for task-specific corpus selection for discriminative feature transformation using HLDA. The first method uses small amount of task-specific corpus to perform discriminative feature transformation. The second method select a subset of the training corpus, based on the task, to perform discriminative feature transformation. The first method uses the existing HMMs to conduct discriminative feature transformation with limited task-specific training data. The second method focuses on selecting specific training data from the existing corpus to conduct discriminative feature transformation with the existing HMMs. The difference between these two methods lies on whether the task-specific corpus is used.The second part of this thesis, features mergence, improves the contextual information of each frame in time domain by cascading the feature of frames. HMMs are then trained with different feature extraction techniques to improve the English speech recognition system. To evaluate the performance of the porposed methods, this thesis uses sentence recognition rate as our performance measure. The experimental result shows that discriminative feature transformation using HLDA has a better performance. Besides, feature mergence also outperforms the baseline acoustic HMMs. Lastly, combining the above two methods achieves the best recognition rate of 97.49% in this research.