Daisy Stanton
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
2024 · arXiv (Cornell University)
Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic …
-
Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis
2019 · arXiv (Cornell University)
Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a coherent framework for understanding the trade-offs between the competing methods. …
-
Tacotron: Towards End-to-End Speech Synthesis
2017
A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design …