ملف الباحث

Navdeep Jaitly

12 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. RNN Approaches to Text Normalization: A Challenge

    2016 · arXiv (Cornell University)

    This paper presents a challenge to the community: given a large corpus of written text aligned to its normalized spoken form, train an RNN to learn the correct normalization function. We present a data set …

  2. Pointer Networks

    2015 · arXiv (Cornell University)

    We introduce a new neural architecture to learn the conditional probability of an output sequence with elements that are discrete tokens corresponding to positions in an input sequence. Such problems cannot be trivially addressed by …

  3. A Simple Way to Initialize Recurrent Networks of Rectified Linear Units

    2015 · arXiv (Cornell University)

    Learning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients. To overcome this difficulty, researchers have developed sophisticated optimization techniques and network architectures. In this paper, we propose a simpler …

  4. Listen, attend and spell: A neural network for large vocabulary conversational speech recognition

    2016

    We present Listen, Attend and Spell (LAS), a neural speech recognizer that transcribes speech utterances directly to characters without pronunciation models, HMMs or other components of traditional speech recognizers. In LAS, the neural network architecture …

  5. Reward Augmented Maximum Likelihood for Neural Structured Prediction

    2016 · arXiv (Cornell University)

    A key problem in structured output prediction is direct optimization of the task reward function that matters for test evaluation. This paper presents a simple and computationally efficient approach to incorporate task reward into a …

  6. Towards Better Decoding and Language Model Integration in Sequence to Sequence Models

    2017

    The recently proposed Sequence-to-Sequence (seq2seq) framework advocates replacing complex data processing pipelines, such as an entire automatic speech recognition system, with a single neural network trained in an end-to-end fashion.In this contribution, we analyse an …

  7. Sequence-to-Sequence Models Can Directly Translate Foreign Speech

    2017

    We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another.The model does not explicitly transcribe the speech into text in the source language, nor does …

  8. Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

    2017 · arXiv (Cornell University)

    This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …

  9. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  10. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks

    2015 · arXiv (Cornell University)

    Recurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning. The current approach to training them consists of maximizing the …

  11. Tacotron: Towards End-to-End Speech Synthesis

    2017

    A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design …

  12. Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions

    2018

    This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …