Arabic Authorship Authentication through Stacked Ensemble Learning
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
Authorship authentication is a linguistic and computational task that aims to determine the actual author of an anonymous or disputed text by analyzing language usage and style. Building upon a significant research gap, this investigation aims to address the Arabic authorship authentication field, which has received less attention due to the structural complexity and syntactic richness of the Arabic language. Despite being the fourth-most used language on the Internet, Arabic lacks good, robust computational methods for authorship analysis. In essence, the main purpose of this research study is to investigate and develop a very robust machine-learning framework that is capable of increasing the precision of Arabic authorship authentication. To this end, a stacking method is proposed as an ensemble tool, composed of several base classifiers, where logistic regression acts as the meta-classifier. A usage-balanced data set of 600 text articles written by ten different writers retrieved from the digital library Alwaraq is established as an experimental corpus. The traditional feature extraction techniques; Bag of Words and TF-IDF are applied, in addition to other hybrid sets including structural, lexical, syntactic, semantic, and stylistic features, to study the whole linguistic variability within Arabic texts. The ensemble stacking models are evaluated by rigorous experimentation, which demonstrates an impressive improvement in classification performance. The proposed model achieves a baseline accuracy of 96.67%, which is improved to 97.5% with the chi-squared feature selection scheme, demonstrating the efficacy of the hybrid features and ensemble structure. The study contributes to computational linguistics and natural language processing via the delivery of a correct and scalable model for Arabic authorship authentication. It also brings out the necessity of expert feature engineering and ensemble learning strategies in overcoming the linguistic complexity of Arabic text classification. The findings have implications in digital forensics, literary critique, and cybersecurity solutions for Arabic-speaking settings.
Publication details
- DOI
- 10.21608/fuje.2025.402811.1118
- OpenAlex
- W4417533681
- Document type
- article
- Language
- EN
- Source
- Fayoum University Journal of Engineering
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.