ملف الباحث

Ron J. Weiss

11 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Sequence-to-Sequence Models Can Directly Translate Foreign Speech

    2017

    We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another.The model does not explicitly transcribe the speech into text in the source language, nor does …

  2. Online and Linear-Time Attention by Enforcing Monotonic Alignments

    2017 · arXiv (Cornell University)

    Recurrent neural network models with an attention mechanism have proven to be extremely effective on a wide variety of sequence-to-sequence problems. However, the fact that soft attention mechanisms perform a pass over the entire input …

  3. Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

    2017 · arXiv (Cornell University)

    This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …

  4. A Spelling Correction Model for End-to-end Speech Recognition

    2019

    Attention-based sequence-to-sequence models for speech recognition jointly train an acoustic model, language model (LM), and alignment mechanism using a single neural network and require only parallel audio-text pairs. Thus, the language model component of the …

  5. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  6. Tacotron: Towards End-to-End Speech Synthesis

    2017

    A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design …

  7. Leveraging Weakly Supervised Data to Improve End-to-end Speech-to-text Translation

    2019

    End-to-end Speech Translation (ST) models have many potential advantages when compared to the cascade of Automatic Speech Recognition (ASR) and text Machine Translation (MT) models, including lowered inference latency and the avoidance of error compounding. …

  8. Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions

    2018

    This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …

  9. Multilingual Speech Recognition with a Single End-to-End Model

    2018

    Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual …

  10. Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning

    2019

    We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish …

  11. Direct Speech-to-Speech Translation with a Sequence-to-Sequence Model

    2019

    We present an attention-based sequence-to-sequence neural network which can directly translate speech from one language into speech in another language, without relying on an intermediate text representation.The network is trained end-to-end, learning to map speech …