Tara N. Sainath
7 papers in the PaperMetrix corpus
Papers by this author
-
Improving Tail Performance of a Deliberation E2E ASR Model Using a Large Text Corpus
2020
End-to-end (E2E) automatic speech recognition (ASR) systems lack the distinct language model (LM) component that characterizes traditional speech systems. While this simplifies the model architecture, it complicates the task of incorporating text-only data into training, …
-
Modular Domain Adaptation for Conformer-Based Streaming ASR
2023 · arXiv (Cornell University)
Speech data from different domains has distinct acoustic and linguistic characteristics. It is common to train a single multidomain model such as a Conformer transducer for speech recognition on a mixture of data from all …
-
Convolutional neural networks for small-footprint keyword spotting
2015
We explore using Convolutional Neural Networks (CNNs) for a small-footprint keyword spotting (KWS) task. CNNs are attractive for KWS since they have been shown to outperform DNNs with far fewer parameters. We consider two different …
-
A Spelling Correction Model for End-to-end Speech Recognition
2019
Attention-based sequence-to-sequence models for speech recognition jointly train an acoustic model, language model (LM), and alignment mechanism using a single neural network and require only parallel audio-text pairs. Thus, the language model component of the …
-
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
2019 · arXiv (Cornell University)
Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …
-
Bytes Are All You Need: End-to-end Multilingual Speech Recognition and Synthesis with Bytes
2019
We present two end-to-end models: Audio-to-Byte (A2B) and Byte-to-Audio (B2A), for multilingual speech recognition and synthesis. Prior work has predominantly used characters, sub-words or words as the unit of choice to model text. These units …
-
Multilingual Speech Recognition with a Single End-to-End Model
2018
Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual …