Abstract
In this study, a singing voice synthesis system is proposed. We improve the naturalness of the synthetic singing voice via the modification of pitch curves. Our goal is to produce a pitch curve similar to that of actual singing voice. We employ two methods for pitch-curve prediction: In the first method, we use support vector machine (SVM) to train a regression model to predict pitch curves. In the second method, we propose a rule-based approach comprising 10 manually-tuned equations for the pitch curves under different conditions. In the second half of the thesis, we discuss the signal processing techniques that are applied to modify pitch, duration and volume. We further solve the problems of ill-articulated pronunciation and discontinuity in the syllable concatenation by using pitch synchronous based crossing fading approach. Moreover, we also create some euphonious effects, such as vibrato and reverberation. Finally, we assess the performance of the proposed methods via pitch curve observation and a listening test experiment. It is verified that the proposed rule-based approach actually is able to make the synthetic singing voices more natural as compared with other traditional singing voice synthesis approaches.