conference-paper

Enhancing Arabic Dialect Classification with Deep Learning Techniques

Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Öz

Arabic dialect classification presents unique challenges due to the linguistic diversity across regions. This study explores deep learning methods to improve classification performance, focusing on underrepresented dialects, especially Saudi variants. We evaluate recurrent neural networks, including BiGRU and BiLSTM, using the SADA dataset—a rich corpus of dialectal Arabic. The approach integrates preprocessing techniques such as normalization, stop-word removal, and tokenization to refine text inputs. Models are assessed using accuracy, precision, recall, and F1-score. Among the architectures, BiLSTM achieved the highest performance with 97% training and 90% test accuracy. To address class imbalance and overfitting, we applied data balancing and regularization strategies, which significantly enhanced generalization. Our findings demonstrate the effectiveness of certain architectures for dialect identification and offer insights for future NLP applications. This work contributes to advancing Arabic dialect processing and provides a baseline for future research in low-resource language classification.

Record transparency

Publication details

DOI
10.1109/esmarta66764.2025.11132273
OpenAlex
W4413679572
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.