Logo image
SPEECH CODING AT RATES BELOW 4.8 KB/S
Dissertation

SPEECH CODING AT RATES BELOW 4.8 KB/S

Kuo, Chih-Chung
Doctor of Philosophy (PHD), 國立清華大學, 電機工程學系
1993

Abstract

語音編碼;低位元率;碼本激發線性預估編碼器 speech coding low bit-rate CELP coder
本篇論文係針對位元率在4.8 kb/s以下之語音編碼的研究。採取的基本架 構是“碼本激發線性預估”(CELP)編碼器。本篇論文共提出了三個新方法 以降低位元率及提高解碼後之語音品質。最後所獲致的成果,是將位元率 降低至3.6 kb/s以下,並保持解碼後之語音品質為通訊品質。第一個新方 法,是針對CELP編碼器中之適應碼本的延遲參數搜尋提出的改進方法。因 為此延遲參數決定解碼後語音的音頻特性,故此延遲參數之變化是否平滑 ,將影響解碼後語音之品質。本法可獲致平滑的延遲變化曲線,並減少延 遲參數搜尋計算量。第二個新方法,是有關CELP編碼器中之線性預估係數 的編碼。將線性預估係數轉成線頻譜對參數之後,利用線頻譜對參數在同 一音框內與連續音框間的關連性,本法設計了一種二維線性預估方法,可 獲得較小變化量之預估差值,從而降低編碼所需位元數。另外,本法也設 計了一種單值量化及二種向量量化的方法,可有效地量化該預估差值。與 直接對線頻譜對參數做單值量化比較,本法約可減少一半所需的位元數 。 CELP編碼器中的適應碼本與隨機碼本,係用來表示語音的激發信號部 份。因此不同類別語音的碼本參數皆有不同的行為特性。本論文提出的第 三個新方法,即是按照長程關連性、週期變化特性、以及語音變化特性, 將每一音框分為無聲、無聲至有聲、有聲、與有聲至無聲四種類別。每一 種類別各有依其特性所設計之碼本參數搜尋與量化的方法。因為分類所根 據的參數,係估測編碼器之碼本參數所產生。故各音框分類的結果,相當 於是其碼本參數特性的估測,因此可使碼本參數的編碼更精確,從而降低 所需的位元率。此外,本法也有減少計算量,與促使適應碼本之延遲參數 變化較平滑的好處。 This dissertation describes a research on techniques for reducing bit-rate below 4.8 kb/s of speech coding with CELP- like coders. Three methods are presented, which remove the redundancy existing in the inter-subframe or inter-frame correlation of the transmitted parameters of a conventional CELP coder. First, a windowed search of the pitch delay which takes advantage of the stable pitch contour in voiced speech segment is designed for constrained adaptive codebook search. This would somewhat lower down the bit-rate and reduce the computation complexity. Second, a new LPC spectral coding method is presented. The LSP parameters which represent the LPC coefficients are viewed as a two dimensional pseudo-stationary signal. A two-dimensional linear prediction technique is used to remove both the intra- and inter-frame redundancy simultaneously. Three special VQ schemes are also designed to quantize the prediction residual efficiently. About half of the bit-rate required to directly quantize LSP parameters is achieved by this new method. Finally, a multi-state CELP coder with a classifier embedded in the adaptive codebook search is shown. This is in fact a further improvement of the windowed search method mentioned above. Each speech frame is classified by a open-loop pitch search and a finite-state machine into one of four states. Then a state-specific adaptive and fixed codebook search is performed to encode each speech subframe with less bits and computations relative to the non-classified coders. All the techniques presented in this dissertation can be integrated into a 3.6 kb/s speech coder that produces near communications quality.

Metrics

1 Record Views

Details

Logo image