Abstract
In electronic music, formant filtering refers to transforming a source signal to sound like vowels. However, existing plugins, such as iZotope VocalSynth 2, and The Orb do not focus on preserving the timbre of the source audio after filtering. In this paper, an algorithm is proposed for replacing a sustained vowel of one singer with another vowel while preserving the timbral characteristics. Here we assume that a timbre source signal (TSS) and a vowel source signal (VSS) are available by recording from singer 1 and 2, respectively. Spectral envelopes below 5k Hz of both signals are extracted first. Then, the ratio of the two envelopes is used for fine-tuning the TSS power spectrum and this process will be repeated iteratively. Upon completion, the output file is meant to preserve the timbre of singer 1 while its vowel identity, including the accent, has been changed to that of singer 2. To evaluate the synthesis quality, sustained vowels sung by 11 singers were recorded, and the proposed algorithm was applied for cross-synthesis. 11 subjects are invited to tell if synthesized timbre resembles singer 1 or singer 2. Results show that 7 subjects successfully chose singer 1 with accuracy > 70%.