RJ Skerry-Ryan
5 papers in the PaperMetrix corpus
Papers by this author
-
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
2024 · arXiv (Cornell University)
Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic …
-
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
2017 · arXiv (Cornell University)
This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …
-
Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis
2019 · arXiv (Cornell University)
Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a coherent framework for understanding the trade-offs between the competing methods. …
-
Tacotron: Towards End-to-End Speech Synthesis
2017
A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design …
-
Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
2019
We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish …