Abstract
The goal of this study it to investigate the proper settings for achieving robust performance of speaker recognition systems under various recording conditions, including microphone, common phones, and mobile phones. For speech feature extraction, we have tried both MFCC (Mel-frequency cepstral coefficients) and wavelet transform coefficients. For classifiers, we have tested GMM (Gaussian mixture models), OGMM (orthogonal GMM), VQOGMM (Vector quantization based OGMM) and ARVM (Auto-regressive vector models). To evaluate the combinations of speech features and speaker classifiers, we have used 3 speech corpora in this study, including TIMIT, NTIMIT, and CTIMIT. The best combination of features and classifiers is not always the same for different corpora. This issue is discussed in details based on empirical result in this thesis.