English Oral Speech Recognition System based on Self-Attention Mechanisms with Bidirectional Long Short-Term Memory
At a glance
- Citations
- 0
- References
- 19
- Comments
- 0
Abstract
Nowadays, English oral speech recognition systems have seen major developments with Deep Learning (DL) techniques, mainly using Recurrent Neural Networks (RNNs) like Long Short-Term Memory (LSTM). However, existing models struggle with long-range dependencies and variability in pronunciation, accents, and background noise and handling diverse speech patterns remain major challenges. To overcome these issues, this research proposed model that combined Self-Attention mechanisms with Bidirectional LSTM (SA-BiLSTM) This procedure captures temporal dependencies while analyzing critical phonetic features, improves recognition accuracy. Initially, data is collected from Google Speech Commands Dataset (GSCD) and preprocessed by using median filter, normalization to improve generalization. Then, feature extraction in done by using Mel-Frequency Cepstral Coefficients (MFCC) and these features are classified by using proposed SA-BiLSTM. From the results, the proposed SA-BiLSTM model attained better outcomes compared to existing Multi-Layer Perceptron-LSTM (MLP-LSTM) in terms of accuracy, precision, recall and F1 Score by achieving 99.02%, 97.31%, 96.76% and 98.32% respectively.
Publication details
- DOI
- 10.1109/iciscn64258.2025.10934181
- OpenAlex
- W4408897200
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.