Abstract
Phone segmentation involves partitioning a continuous speech signal into discrete phone units. It is often required in some areas of speech processing, such as acoustic-phonetic analysis, speech recognition, speaker recognition, speech synthesis, and annotations of speech corpus. Manual phone segmentation is time consuming, and its result may be inconsistent because of the subjective criteria of different transcribers. Therefore a method of automatic phone segmentation is desirable. A typical approach is to align the speech signal to its phone transcripts in an utterance. The forced alignment based on hidden Markov model is a way to locate phone boundaries when the phone transcripts of the target utterance are available. This supervised method usually obtains high accuracy. However, the training speech signal and their transcripts are unavailable in some applications. Hence, unsupervised methods are used. If there is no linguistic knowledge (such as, orthographic or phonetic transcripts) of given speech data, phone segmentation is performed in blind method. However, this approach is difficult to obtain a high accuracy. Obtaining a high level of accuracy by using the blind method is challenging. This dissertation addresses the problem of blind phone segmentation. The band energies of speech signals are calculated for feature extraction. Four methods for blind phone segmentation are proposed. They are based on (1)Delta spectral function, (2)Band-energy tracing technique, (3)Gaussian function, and (4)Legendre polynomial approximation. English speech corpus, TIMIT, was examined. Experimental results showed that the proposed methods were more accurate than previous methods. For the method using BE tracing technique, Chinese speech corpus, TCC300, was also evaluated to reveal the language-independent problems. Noise influences were investigated in the methods using Gaussian function and Legendre polynomial approximation.