Yatharth Saraf
4 papers in the PaperMetrix corpus
Papers by this author
-
Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
2020 · arXiv (Cornell University)
In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state-of-the-art results. We then show that using wordpieces as modeling units combined with …
-
Improving RNN Transducer Based ASR with Auxiliary Tasks
2021
End-to-end automatic speech recognition (ASR) models with a single neural network have recently demonstrated state-of-the-art results compared to conventional hybrid speech recognizers. Specifically, recurrent neural network transducer (RNN-T) has shown competitive ASR performance on various …
-
Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings
2021 · arXiv (Cornell University)
Speech recognition models often obtain degraded performance when tested on speech with unseen accents. Domain-adversarial training (DAT) and multi-task learning (MTL) are two common approaches for building accent-robust ASR models. ASR models using accent embeddings …
-
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
2022 · Interspeech 2022
This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0.We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in …