Nepali Speech Recognition using CNN and Sequence Models
At a glance
- Citations
- 2
- References
- 12
- Comments
- 0
Öz
Speech-To-Text(STT) is becoming an important way that humans interact with computers. STT systems are quite popular nowadays and are widely used for many applications and are available in many languages. However, for Nepali language, such systems are not readily available. In this paper, we implement several models for STT, based on Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN, GRU). We also compare these different models and their combinations. The models we present transcribe audio character-by-character. The predictions are made and the highest probability transcription is selected using a beam search decoder. For training, a dataset from OpenSLR was used. The best trained model (CNN-RNN) has an character error rate of 23.72% on the validation data.
Publication details
- DOI
- 10.1109/icmlant50963.2020.9355707
- OpenAlex
- W3132829453
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.