Real-Time One-Pass Decoder for Speech Recognition Using LSTM Language Models
At a glance
- الاستشهادات
- 23
- المراجع
- 20
- Comments
- 0
Abstract
Recurrent Neural Networks, in particular Long-Short TermMemory (LSTM) networks, are widely used in Automatic Speech Recognition for language modelling during decoding,usually as a mechanism for rescoring hypothesis. This paperproposes a new architecture to perform real-time one-pass de-coding using LSTM language models. To make decoding ef-ficient, the estimation of look-ahead scores was accelerated byprecomputing static look-ahead tables. These static tables wereprecomputed from a prunedn-gram model, reducing drasti-cally the computational cost during decoding. Additionally,the LSTM language model evaluation was efficiently performedusing Variance Regularization along with a strategy of lazyevaluation. The proposed one-pass decoder architecture wasevaluated on the well-known LibriSpeech and TED-LIUMv3datasets. Results showed that the proposed algorithm obtainsvery competitive WERs with 0.6 RTFs. Finally, our one-passdecoder is compared with a decoupled two-pass decoder.
Publication details
- DOI
- 10.21437/interspeech.2019-2798
- OpenAlex
- W2972528057
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.