conference-paper

Nepali Speech Recognition using CNN and Sequence Models

Research footprint

At a glance

Citations
2
References
12
Comments
0
Paper overview

Öz

Speech-To-Text(STT) is becoming an important way that humans interact with computers. STT systems are quite popular nowadays and are widely used for many applications and are available in many languages. However, for Nepali language, such systems are not readily available. In this paper, we implement several models for STT, based on Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN, GRU). We also compare these different models and their combinations. The models we present transcribe audio character-by-character. The predictions are made and the highest probability transcription is selected using a beam search decoder. For training, a dataset from OpenSLR was used. The best trained model (CNN-RNN) has an character error rate of 23.72% on the validation data.

Record transparency

Publication details

DOI
10.1109/icmlant50963.2020.9355707
OpenAlex
W3132829453
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.