Abstract
當語音辨認系統應用在電話環境上時,由於訓練與測試語音環境的不同,常導致辨認效果的衰減,在電話環境上的失真來源包括有雜訊、通道和語者的失真,本論文提出一系列強健性演算法做三種失真來源的補償,以提昇辨認效果。在隱藏式馬可夫模型為主的語音辨認實驗裡,本論文所提出的方法均能成功的克服電話環境下的失真問題。 論文首先分析雜訊效應對語音逆頻譜向量及隱藏式馬可夫模型參數的影響。由於逆頻譜向量受雜訊的干擾會畏縮,將隱藏式馬可夫模型參數的平均值向量用最佳畏縮因子做調整所發展出來的投射性相似測量,對雜訊的干擾具有強健性的效果,本論文延伸此研究,進一步補償模型參數的變異數畏縮以及平均值的調整偏差,補償的因子是由一組調整函數所獲得,實驗證明,使用本方法的辨認率有明顯的提昇。 為了克服電話語音的通道效應,本論文發展出一種通道效應消除法,本方法是先量化一些電話倒通道模型的逆頻譜向量,以訓練出一組參考濾波器,而通道效應消除濾波器的逆頻譜向量就是由這組參考濾波器的逆頻譜向量所線性組合而成的,其組合係數的求法是根據電話語音通過參考濾波器的累積觀察機率所求得,此方法可以有效消除電話語音的通道效應。其次,本論文提出兩種轉換式的調整方法以調整隱藏式馬可夫模型參數,使調整過的模型參數能夠較接近於測試時的電話環境,這兩種轉換式調整法分別是偏差轉換及仿射轉換,我們使用有考慮事前統計特性的最佳事後機率法則做轉換參數的估測,在我們的實驗評估裡發現,使用最佳事後機率法則的調整方法比使用最佳相似法則的效果好,而且仿射轉換的精確度優於偏差轉換。此外,本論文也提出一種音相關通道補償法,此方法是利用一些調整語句將原始的隱藏式馬可夫模型參數調整到新的通道環境下使用,調整的方法是將模型參數結合上其對應的音相關 通道補償向量,為了改善調整效果,我們提出兩種延伸技術,第一種是利用向量量化法將補償向量的精確度提高,第二種是利用外差法將補償向量做線性外差,這兩種技術都已成功的應用在電話語音辨認及語者調適上。 另外,我們也提出一種混合式演算法,將非特定語者之隱藏式馬可夫模型調整到新的語者特性上,本方法是結合三種調整技術,首先,將不同群組的模型參數用相對應的轉換函數做轉換,然後將轉換過的模型參數做最佳事後機率調整,最後,在最佳事後機率調整裡未調整到的模型參數用轉移向量外差法做進一步的調整,實驗發現使用本方法可以同時達到這三種調整技術的優點,在不同長度的調整語句下都比其他調整方法的效果好。When the speech recognition system is operated undertelephone networks, the acoustic mismatch between training andtesting environments always causes the performance degradation.The mismatch sources in telephone environments areattributed tothe ambient noise, the channel effect and the variation amongspeakers. This dissertation describes a number of robustalgorithms which improve the recognition performance bycompensating these three mismatch factors. In the experiments ofhidden Markov model (HMM) based speech recognition, the proposedmethods can successfully overcome the mismatch problems intelephone environments. The noise effect on speech cepstralvector and its associated HMM acoustic parameters is firstinvestigated. Due to the shrinkage of cepstral vector in noisyenvironment, the projection-based likelihood measure which usesan optimalequalization factor for adapting the cepstral meanvector of HMM parameters is robust to noise contamination. Thisdissertation extends this measure by further compensating theshrinkage of covariance matrix and the bias of mean vector. Thecompensation factors are obtained from a set of adaptationfunctions. Using this method, the recognition accuracy can beremarkably improved. To overcome the channel effect intelephone speech, a channel-effect-cancellation method isdeveloped. This approach is to estimate a channel-effect-cancellation filter by the convex combination of severalreference filters. The reference filters, represented incepstrum, are generated by clustering the cepstra of inversetelephone channels. The convex combination coefficients arecalculated by the accumulated observation probabilities when thetesting utterance passes through the reference filters. Usingthis method,the channel effect can be mostly canceled. Next,this dissertation presents two transformation-based adaptationapproaches for adapting the HMM parameters so that the adaptedHMM parameters are acoustically close to the telephoneenvironment. The bias and the affine transformations areexamined. We apply the maximum a posteriori (MAP) estimationtechnique which incorporates the prior knowledge into thetransformation for estimating the transformation parameters.Inour evaluation, the transformation-based adaptation using theMAP estimationoutperforms that using the maximum likelihood (ML)estimation. The affine transformation is also demonstrated to besuperior to the bias transformation. Furthermore, a phone-dependent channel compensation (PDCC) technique is proposed foradapting the HMM parameters to a new channel environment byusing some adaptation data. The adaptation of HMM parameters iscompleted by incorporating the corresponding PDCC vectors. Toimprove the performance, two extended PDCC techniques arepresented. One is based on the refinement of PDCC using vectorquantization. The other is based on the interpolation ofcompensation vectors. This method is carried out and shown to beeffective in telephone speech recognition as well as speakeradaptation. In addition, we also propose a hybrid algorithmfor adapting the HMM parameters to a new speaker. This algorithmis constructed by iteratively and alternately combining threeadaptation techniques. First, the clusters of HMM parameters arelocally transformed through a group of transformation functions.Then, the transformed HMM parameters are globally smoothed viathe MAP adaptation. Within the MAP adaptation, the parameters ofunseen units in adaptation data are further adapted by applyingthe transfer vector interpolation scheme. Using this algorithm,the advantages of these three adaptation techniques can besimultaneously captured. The resulting performance isconsistently better than other methods for almost any practicalamount of adaptation data.