Reinforcement Learning for Speech Recognition using Recurrent Neural Networks
At a glance
- الاستشهادات
- 5
- المراجع
- 8
- Comments
- 0
Abstract
This work describes a voice recognition system that does not need an intermediate phonetic representation to convert audio input to text. The system is based on a mix of the the Connectionist Temporal Classification goal function and deep bidirectional LSTM recurrent neural network architecture . A new method is proposed in which the network is taught to reduce the likelihood of an arbitrary transcription loss function being encountered. without the aid of any lexicons or models, this allows for a direct optimization of WER. The system has a WER (word error rate) of 22 percent, 20 percent with simply a lexicon of authorized terms, 9 percent using a trigram language model. The error rate drops to 7 percent when the network is used in conjunction with a baseline system.
Publication details
- DOI
- 10.1109/asiancon55314.2022.9908930
- OpenAlex
- W4304208085
- Document type
- conference-paper
- Language
- EN
- Source
- 2022 2nd Asian Conference on Innovation in Technology (ASIANCON)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.