conference-paper
Speaker and language factorization in DNN-based TTS synthesis
Research footprint
At a glance
- Citations
- 35
- References
- 14
- Comments
- 0
Paper overview
Öz
We have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, where speaker-specific layers are used for multi-speaker modelling, and shared layers and language-specific layers are employed for multi-language, linguistic feature transformation. Experimental results on a speech corpus of multiple speakers in both Mandarin and English show that the proposed factorized DNN can not only achieve a similar voice quality as that of a multi-speaker DNN, but also perform polyglot synthesis with a monolingual speaker's voice.
Record transparency
Publication details
- DOI
- 10.1109/icassp.2016.7472737
- OpenAlex
- W2401698713
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.