conference-paper

Automated Child Voice Generation: Methodology and Implementation

Research footprint

At a glance

الاستشهادات
2
المراجع
23
Comments
0
Paper overview

Abstract

Significant progress has been made in the development of text-to-speech (TTS) models; however, synthesizing child speech remains a challenging task. Limited research has been conducted on this topic due to the lack of child speech datasets and the inherent difficulties in constructing such datasets. Children's speech is often less clear and exhibits significant variations in terms of volume, pitch, and rhythm. In this study, we explore the use of two different vocoders for synthesizing conversational multi-speaker child speech: the WORLD vocoder based on statistical parametric speech synthesis (SPSS) and the neural vocoders based on Parallel WaveGAN and AutoVocoder. Initially, we trained the AutoVocoder on a dataset of female adult speech. Subsequently, we investigated the effectiveness of fine-tuning and adapting these vocoders to capture the unique characteristics of child speech, while mitigating the need for extensive child speech datasets. Experimental results demonstrated that the AutoVocoder outperformed other vocoders in terms of clarity when synthesizing conversational multi-speaker child speech. Despite the challenges posed by the MyST child datasets used in this study, which included non-phonetic noise and indiscernible speech, the AutoVocoder significantly improved the quality and clarity of the ground truth in the context of conversational multi-speaker child speech synthesis. Both objective and subjective evaluations indicated that the original and synthesized speech by the AutoVocoder were very similar to each other.

Record transparency

Publication details

DOI
10.1109/sped59241.2023.10314889
OpenAlex
W4388692945
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.