article

FGP-GAN: Fine-Grained Perception Integrated Generative Adversarial Network for Expressive Mandarin Singing Voice Synthesis

  • IEEE Transactions on Consumer Electronics
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
3
References
45
Comments
0
Paper overview

Abstract

Though singing voice synthesis (SVS) has been explored by recent works, pitch over-smoothing and spectral blurring are still unresolved, resulting in lack of expressiveness of the singing voice. To tackle the afore-mentioned problems, an SVS model that adopts fine-grained perception modeling in the generative adversarial framework is proposed in this work. Specifically, the forward difference of fundamental frequency is also modeled in addition to the conventional fundamental frequency in the generator, resulting in more accurate vocal fundamental frequencies. Moreover, the fine-grained perception module is utilized to restrict the generator to focus more on detailed information. Then, we adopt the WGAN discriminator to optimize the SVS model such that the distribution of the generated singing spectrogram effectively approximates the actual one. Finally, the spectrogram was fed into our modified vocoder, resulting in more expressive and natural singing voice. The effectiveness of the proposed approach is verified on a professional Mandarin Chinese corpus. Experimental results demonstrate that the proposed approach can obtain more accurate fundamental frequencies, clearer spectrograms, more natural singing voice and higher mean opinion score (MOS), comparing with several state-of-the-art approaches.

Record transparency

Publication details

DOI
10.1109/tce.2024.3412053
OpenAlex
W4399527496
Document type
article
Language
EN
Source
IEEE Transactions on Consumer Electronics
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.