Navdeep Jaitly
12 papers in the PaperMetrix corpus
Papers by this author
-
RNN Approaches to Text Normalization: A Challenge
2016 · arXiv (Cornell University)
This paper presents a challenge to the community: given a large corpus of written text aligned to its normalized spoken form, train an RNN to learn the correct normalization function. We present a data set …
-
Pointer Networks
2015 · arXiv (Cornell University)
We introduce a new neural architecture to learn the conditional probability of an output sequence with elements that are discrete tokens corresponding to positions in an input sequence. Such problems cannot be trivially addressed by …
-
A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
2015 · arXiv (Cornell University)
Learning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients. To overcome this difficulty, researchers have developed sophisticated optimization techniques and network architectures. In this paper, we propose a simpler …
-
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
2016
We present Listen, Attend and Spell (LAS), a neural speech recognizer that transcribes speech utterances directly to characters without pronunciation models, HMMs or other components of traditional speech recognizers. In LAS, the neural network architecture …
-
Reward Augmented Maximum Likelihood for Neural Structured Prediction
2016 · arXiv (Cornell University)
A key problem in structured output prediction is direct optimization of the task reward function that matters for test evaluation. This paper presents a simple and computationally efficient approach to incorporate task reward into a …
-
Towards Better Decoding and Language Model Integration in Sequence to Sequence Models
2017
The recently proposed Sequence-to-Sequence (seq2seq) framework advocates replacing complex data processing pipelines, such as an entire automatic speech recognition system, with a single neural network trained in an end-to-end fashion.In this contribution, we analyse an …
-
Sequence-to-Sequence Models Can Directly Translate Foreign Speech
2017
We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another.The model does not explicitly transcribe the speech into text in the source language, nor does …
-
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
2017 · arXiv (Cornell University)
This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …
-
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
2019 · arXiv (Cornell University)
Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …
-
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
2015 · arXiv (Cornell University)
Recurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning. The current approach to training them consists of maximizing the …
-
Tacotron: Towards End-to-End Speech Synthesis
2017
A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design …
-
Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions
2018
This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …