Automated Child Voice Generation: Methodology and Implementation
At a glance
- الاستشهادات
- 2
- المراجع
- 23
- Comments
- 0
Abstract
Significant progress has been made in the development of text-to-speech (TTS) models; however, synthesizing child speech remains a challenging task. Limited research has been conducted on this topic due to the lack of child speech datasets and the inherent difficulties in constructing such datasets. Children's speech is often less clear and exhibits significant variations in terms of volume, pitch, and rhythm. In this study, we explore the use of two different vocoders for synthesizing conversational multi-speaker child speech: the WORLD vocoder based on statistical parametric speech synthesis (SPSS) and the neural vocoders based on Parallel WaveGAN and AutoVocoder. Initially, we trained the AutoVocoder on a dataset of female adult speech. Subsequently, we investigated the effectiveness of fine-tuning and adapting these vocoders to capture the unique characteristics of child speech, while mitigating the need for extensive child speech datasets. Experimental results demonstrated that the AutoVocoder outperformed other vocoders in terms of clarity when synthesizing conversational multi-speaker child speech. Despite the challenges posed by the MyST child datasets used in this study, which included non-phonetic noise and indiscernible speech, the AutoVocoder significantly improved the quality and clarity of the ground truth in the context of conversational multi-speaker child speech synthesis. Both objective and subjective evaluations indicated that the original and synthesized speech by the AutoVocoder were very similar to each other.
Publication details
- DOI
- 10.1109/sped59241.2023.10314889
- OpenAlex
- W4388692945
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.