Logo image
Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN
Conference paper

Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN

Yin-Ping Cho, Yu Tsao, Hsin-Min Wang and Yi-Wen Liu
Proceedings of 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2022, pp.1956-1963
2022

Abstract

Computer Networks and Communications Information Systems Signal Processing
Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this work adopts the acoustic model-neural vocoder architecture established for high-quality speech and singing voice synthesis. Specifically, this work aims to pursue a higher level of expres-siveness in synthesized voices by combining the diffusion de-noising probabilistic model (DDPM) and Wasserstein generative adversarial network (WGAN) to construct the backbone of the acoustic model. On top of the proposed acoustic model, a HiFi-GAN neural vocoder is adopted with integrated fine-tuning to ensure optimal synthesis quality for the resulting end-to-end SVS system. This end-to-end system was evaluated with the multi-singer Mpop600 Mandarin singing voice dataset. In the exper-iments, the proposed system exhibits improvements over previ-ous landmark counterparts in terms of musical expressiveness and high-frequency acoustic details. Moreover, the adversarial acoustic model converged stably without the need to enforce reconstruction objectives, indicating the convergence stability of the proposed DDPM and WGAN combined architecture over alternative GAN-based SVS systems.11Evaluation audio samples can be found at: https://yinping-cho.github.io/ diffwgansvs.github.io/

Metrics

1 Record Views

Details

Logo image