Abstract
This thesis presents a concatenation-based singing voice synthesis system for Chinese songs. The system takes melody and lyrics information from a KAR file (a variant of MIDI with lyrics information) and employs text-to-speech techniques to synthesis Chinese songs. According to the lyrics information, the system first selects suitable syllable clips from a pre-recorded collection of all 411 syllables in Mandarin Chinese. Then the system performs necessary pitch shift on the syllabic clips using various methods including PSOLA (Pitch Synchronous Overlap and Add), Cross-fading, Resample, Residual Signal with PSOLA, etc. Time-scale modification is then achieved by a linear mapping and duplication of pitch-mark-justified waveforms. With correct pitch and duration, the resulting vocal clips are then concatenated to form a complete vocal rendition of the song. To make natural-sounding singing voice, the system has to employ several post-processing methods including the addition of coarticulation and vibrato, and the use of energy normalization. Potential applications and future research directions are also covered in the thesis.