Researcher profile

RJ Skerry-Ryan

5 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech

    2024 · arXiv (Cornell University)

    Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic …

  2. Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

    2017 · arXiv (Cornell University)

    This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …

  3. Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis

    2019 · arXiv (Cornell University)

    Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a coherent framework for understanding the trade-offs between the competing methods. …

  4. Tacotron: Towards End-to-End Speech Synthesis

    2017

    A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design …

  5. Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning

    2019

    We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish …