A Pitch-Controlled End-to-End Voice Conversion System for Brazilian Portuguese
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
Speech conversion is a technique that modifies the identity of the voice in a speech signal without changing the spoken content. Accurate pitch conversion is a requirement the best speech conversion systems must address, as this characteristic is essential to the correct identification of the target speaker. This work proposes a pitch-controlled end-to-end voice conversion model that combines state-of-the-art ideas from both speaking and singing voice conversion with a novel cost function to ensure artifact-free pitch tracking. The model is trained in Brazilian Portuguese, overcoming the lack of high-quality data by improving a large but flawed dataset with filtering operation. Our model mostly outperforms other popular open source models in both listening tests and objective measurements. In particular, on a 5-point MOS, we obtained the highest speaker similarity score (4.05), and a naturalness score of 3.48, second only to a system whose similarity score was 2.62.
Publication details
- DOI
- 10.14209/jcis.2024.13
- OpenAlex
- W4401529376
- Document type
- article
- Language
- EN
- Source
- Journal of Communication and Information Systems
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.