George Saon
3 papers in the PaperMetrix corpus
Papers by this author
-
Accelerating deep neural network learning for speech recognition on a cluster of GPUs
2017
We train deep neural networks to solve the acoustic modeling problem for large-vocabulary continuous speech recognition. We employ distributed processing using a cluster of GPUs. On modern GPUs, the sequential implementation takes over a day …
-
RNN Transducer Models for Spoken Language Understanding
2021
We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, …
-
Exploring the limits of decoder-only models trained on public speech recognition corpora
2024 · arXiv (Cornell University)
The emergence of industrial-scale speech recognition (ASR) models such as Whisper and USM, trained on 1M hours of weakly labelled and 12M hours of audio only proprietary data respectively, has led to a stronger need …