Geoffrey Zweig
6 papers in the PaperMetrix corpus
Papers by this author
-
The Microsoft 2016 Conversational Speech Recognition System
2016 · arXiv (Cornell University)
We describe the 2017 version of Microsoft's conversational speech recognition system, in which we update our 2016 system with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art …
-
Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
2020 · arXiv (Cornell University)
In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state-of-the-art results. We then show that using wordpieces as modeling units combined with …
-
Improving RNN Transducer Based ASR with Auxiliary Tasks
2021
End-to-end automatic speech recognition (ASR) models with a single neural network have recently demonstrated state-of-the-art results compared to conventional hybrid speech recognizers. Specifically, recurrent neural network transducer (RNN-T) has shown competitive ASR performance on various …
-
Sequence-to-Sequence Neural Net Models for Grapheme-to-Phoneme Conversion
2015 · arXiv (Cornell University)
Sequence-to-sequence translation methods based on generation with a side-conditioned language model have recently shown promising results in several tasks. In machine translation, models conditioned on source side words have been used to produce target-language text, …
-
Hybrid Code Networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning
2017
End-to-end learning of recurrent neural networks (RNNs) is an attractive solution for dialog systems; however, current techniques are data-intensive and require thousands of dialogs to learn simple behaviors.
-
From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition
2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
There is an implicit assumption that traditional hybrid approaches for automatic speech recognition (ASR) cannot directly model graphemes and need to rely on phonetic lexicons to get competitive performance, especially on English which has poor …