Eric Battenberg
3 papers in the PaperMetrix corpus
Papers by this author
-
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
2024 · arXiv (Cornell University)
Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic …
-
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
2015 · arXiv (Cornell University)
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning …
-
Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis
2019 · arXiv (Cornell University)
Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a coherent framework for understanding the trade-offs between the competing methods. …