conference-paper

Excitation-by-SampleRNN Model for Text-to-Speech

  • 2019 34th International Technical Conference on Circuits/Systems, Computers and Communications (ITC-CSCC)
Research footprint

At a glance

Citations
1
References
21
Comments
0
Paper overview

Abstract

In this paper, we propose a neural vocoder-based text-to-speech (TTS) system that effectively utilizes a source-filter modeling framework. Although neural vocoder algorithms such as SampleRNN and WaveNet are well-known to generate high-quality speech, its generation speed is too slow to be used for real-world applications. By decomposing a speech signal into spectral and excitation components based on a source-filter framework, we train those two components separately, i.e. training the spectrum or acoustic parameters with a long short-term memory model and the excitation component with a SampleRNN-based generative model. Unlike the conventional generative model that needs to represent the complicated probabilistic distribution of speech waveform, the proposed approach needs to generate only the glottal movement of human production mechanism. Therefore, it is possible to obtain high-quality speech signals using a small-size of the pitch interval-oriented SampleRNN network. The objective and subjective test results confirm the superiority of the proposed system over a glottal modeling-based parametric and original SampleRNN-based speech synthesis systems.

Record transparency

Publication details

DOI
10.1109/itc-cscc.2019.8793459
OpenAlex
W2968397919
Document type
conference-paper
Language
EN
Source
2019 34th International Technical Conference on Circuits/Systems, Computers and Communications (ITC-CSCC)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.