conference-paper

English Oral Speech Recognition System based on Self-Attention Mechanisms with Bidirectional Long Short-Term Memory

Research footprint

At a glance

الاستشهادات
0
المراجع
19
Comments
0
Paper overview

Abstract

Nowadays, English oral speech recognition systems have seen major developments with Deep Learning (DL) techniques, mainly using Recurrent Neural Networks (RNNs) like Long Short-Term Memory (LSTM). However, existing models struggle with long-range dependencies and variability in pronunciation, accents, and background noise and handling diverse speech patterns remain major challenges. To overcome these issues, this research proposed model that combined Self-Attention mechanisms with Bidirectional LSTM (SA-BiLSTM) This procedure captures temporal dependencies while analyzing critical phonetic features, improves recognition accuracy. Initially, data is collected from Google Speech Commands Dataset (GSCD) and preprocessed by using median filter, normalization to improve generalization. Then, feature extraction in done by using Mel-Frequency Cepstral Coefficients (MFCC) and these features are classified by using proposed SA-BiLSTM. From the results, the proposed SA-BiLSTM model attained better outcomes compared to existing Multi-Layer Perceptron-LSTM (MLP-LSTM) in terms of accuracy, precision, recall and F1 Score by achieving 99.02%, 97.31%, 96.76% and 98.32% respectively.

Record transparency

Publication details

DOI
10.1109/iciscn64258.2025.10934181
OpenAlex
W4408897200
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.