Brian Kingsbury
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Accelerating deep neural network learning for speech recognition on a cluster of GPUs
2017
We train deep neural networks to solve the acoustic modeling problem for large-vocabulary continuous speech recognition. We employ distributed processing using a cluster of GPUs. On modern GPUs, the sequential implementation takes over a day …
-
Beyond Backprop: Alternating Minimization with co-Activation Memory.
2018 · arXiv (Cornell University)
Despite significant recent advances in deep neural networks, training them remains a challenge due to the highly non-convex nature of the objective function. State-of-the-art methods rely on error backpropagation, which suffers from several well-known issues, …
-
Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems
2020 · arXiv (Cornell University)
Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with …
-
RNN Transducer Models for Spoken Language Understanding
2021
We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, …
-
Exploring the limits of decoder-only models trained on public speech recognition corpora
2024 · arXiv (Cornell University)
The emergence of industrial-scale speech recognition (ASR) models such as Whisper and USM, trained on 1M hours of weakly labelled and 12M hours of audio only proprietary data respectively, has led to a stronger need …