article Open access

Text-to-speech with linear spectrogram prediction for quality and speed improvement

  • Phonetics and Speech Sciences
Research footprint

At a glance

Citations
0
References
21
Comments
0
Paper overview

Abstract

Most neural-network-based speech synthesis models utilize neural vocoders to convert mel-scaled spectrograms into high-quality, human-like voices. However, neural vocoders combined with mel-scaled spectrogram prediction models demand considerable computer memory and time during the training phase and are subject to slow inference speeds in an environment where GPU is not used. This problem does not arise in linear spectrogram prediction models, as they do not use neural vocoders, but these models suffer from low voice quality. As a solution, this paper proposes a Tacotron 2 and Transformer-based linear spectrogram prediction model that produces high-quality speech and does not use neural vocoders. Experiments suggest that this model can serve as the foundation of a high-quality text-to-speech model with fast inference speed.

Record transparency

Publication details

DOI
10.13064/ksss.2021.13.3.071
OpenAlex
W3207567141
Document type
article
Language
EN
Source
Phonetics and Speech Sciences
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.