Researcher profile

Tara N. Sainath

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Improving Tail Performance of a Deliberation E2E ASR Model Using a Large Text Corpus

    2020

    End-to-end (E2E) automatic speech recognition (ASR) systems lack the distinct language model (LM) component that characterizes traditional speech systems. While this simplifies the model architecture, it complicates the task of incorporating text-only data into training, …

  2. Modular Domain Adaptation for Conformer-Based Streaming ASR

    2023 · arXiv (Cornell University)

    Speech data from different domains has distinct acoustic and linguistic characteristics. It is common to train a single multidomain model such as a Conformer transducer for speech recognition on a mixture of data from all …

  3. Convolutional neural networks for small-footprint keyword spotting

    2015

    We explore using Convolutional Neural Networks (CNNs) for a small-footprint keyword spotting (KWS) task. CNNs are attractive for KWS since they have been shown to outperform DNNs with far fewer parameters. We consider two different …

  4. A Spelling Correction Model for End-to-end Speech Recognition

    2019

    Attention-based sequence-to-sequence models for speech recognition jointly train an acoustic model, language model (LM), and alignment mechanism using a single neural network and require only parallel audio-text pairs. Thus, the language model component of the …

  5. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  6. Bytes Are All You Need: End-to-end Multilingual Speech Recognition and Synthesis with Bytes

    2019

    We present two end-to-end models: Audio-to-Byte (A2B) and Byte-to-Audio (B2A), for multilingual speech recognition and synthesis. Prior work has predominantly used characters, sub-words or words as the unit of choice to model text. These units …

  7. Multilingual Speech Recognition with a Single End-to-End Model

    2018

    Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual …