Logo image
國語子音辨認的研究
Thesis

國語子音辨認的研究

劉利誠
Masters, National Tsing Hua University
1990

Abstract

隱藏式馬可夫模式特徵向量貝氏分類器神經網路分類器最大可能 HIDDEN-MAKOV-MODELFEATURE-CECTORMAXIMUM-LIKELIHOOD
本文提出一個新的方法,作為隱藏式馬可夫模式參數的訓練。代表特徵向量空間之所有的特徵向量的一組訓練資料被區分為幾個區域,每一個區域中的特徵向量假設是由某一個特徵機率分布隨機產生的;而每一個特徵機率分布對應一個特定的狀態。這樣的作法相當於找出語音訊號所有可能處在的狀態及其特徵向量的機率分布。所有的特徵向量之機率分布的求得有得兩種方法:(一)VQ+ML定義一個不相似度之量測,採用向量量化之方法將特徵向量空間區分為幾個區域,之後再採用最大可能的估算法,把每一個區域中的特徵向量的機率分布估算出來。(二)最大可能的重複演算法不需定義于相似度之量測,先假設起始的一組特徵向量的機率分布,再重複使用最大可能的決策及最大可能的估算,來獲得最後收斂的一組特徵向量的機率分布。當有的特徵向量的機率分布獲得之後,再採用一步的估算法,來獲得馬可夫模式的參數。這方法應用在辨認52個國語音節上,這52個音節是21個聲母後面接/i,a,u,1/ 的所有存在的音節。和傳統之訓練方法比較,所提出的方法具有下列的好處:(一)不需先假設模式中狀態的數目。(二)要辨認新的語者所講到語音,只需重新訓練模式中狀態轉移機率及起始狀態的機率分布,即可達到良好的辨認結果。由於無聲非送氣的塞音,其音長短音強弱,故較不易辨認,本文提出一組特徵來對這些塞音作有效的分類;所採用的分類器有貝式分類器及類神經網路分類器,本文對這兩個分類器作了些比較,並對類神經網路分類器也作了一些特性的分析。///////ABSTRACTHidden Markov modeling has been widely used in the recognition of speechpatterns. For each speech-pattern class, there is a finite-state Markovmodel for the pattern-generating mechanism. A set of training utterancesis required for training the models so that the model parameters can beobtained. A well-known training method for obtaining the model parametersis Baum-Welch reestimation algorithm. This algorithm requires iterativeadjustment of the model starting from an initial estimate. The number ofstates and the model structure of each HMM is empirically determined.In this dissertation, We propose an alternative method for training themodels. The feature vectors obtained from a set oa training data aregrouped into a number of clusters. The feature vectors in a cluster areassumed to be the random samples from one of the feature distributions.Each distribution is associated with a certain state. The maximum-likelihood estimation is applied to obtain the distribution of thefeatuer vectors in each cluster. In the conventional vector quantization,the criterion of finding code-words is to minimize the total distancesbetween training samples and its quantized code-words. In this study, Weuse maximum-likelihood as the clustering criterion, For M clusters. thereare M feature distributios, bi(o), i=1, 2,..., M, which are obtained bymaximizing the following jount(圖表省略)where Q is the total number of training data. Each distribution isconsidered to be associated with a state. Then, an utterance can berepresented by a sequence of states, and the speech signal is modeled by aMarkov chain with given state-transition probabilities. In thisdissertation, the state transition probabilities and the initial-statedistribution probabilities are determined by one-step maximum-likelihoodestimation technique. This gives much simpler training procedure ascomparing to Baum-Welch reestimation algorithm. The proposed method isapplied to the recognition of 52 confusion CVsyllables. The recognitionresults show that the proposed method is superior to continuous HMM. Alarge part of the recognition errors, for some speakers, are those causedby the confusions among voiceless unaspirated stops. Short duration andlow intensity make the recognition of these stops difficult.This dissertation also presents a method of feature extraction for theautomatic recognition of voiceless unaspirated stop consonants. Thefeatures are derived from the spectrographic acoustic patterns ofsyllable-initial voiceless unaspirated stops /p, t, k/, which includethe burst spectrum, the formant transition, and the voice onset time. Anormalization process for the second and the third formants at thevoice-onset is proposed. Based on these derived features, Bayesclassifiers and a layered neural ner are applied to classify the places ofarticulation of these stop consonants. The experiments show that thederived features are robust and efficient for apeaker-independent speechrecognition, and the neural net is a preferable choice in theclassification of these stops in multiple contexts.

Metrics

1 Record Views

Details

Logo image