Performance Comparison of SBERT and FastText in Social Media Analysis
At a glance
- الاستشهادات
- 0
- المراجع
- 11
- Comments
- 0
Abstract
Social media platforms such as Discord often create a number of hate speech content related to ethnicity, religion, race, and inter-group relations (SARA), posing challenges to user comfort and community harmony. This study investigates the effectiveness of SBERT and FastText embeddings in combination with Naive Bayes and Support Vector Machine (SVM) classifiers for text classification tasks and the VADER method for sentiment analysis, using Indonesian-language Discord data. This work demonstrates that SVM outperforms Naive Bayes across all evaluation metrics, achieving higher AUC (86.3% vs. 79.3%), accuracy (83% vs. 71.8%), F1-Score (82.1% vs. 72.9%), Precision (82.8% vs. 79.1%), Recall (83% vs. 71.8%), and MCC (58.2% vs. 47.1%). The SVM outperforms high-dimensional data, and its robust margin-based decision making contributes to better performance in detecting SARA content. These findings underscore the importance of selecting appropriate embedding models and classifiers to optimize hate speech detection and sentiment analysis in informal and diverse linguistic contexts.
Publication details
- DOI
- 10.1109/icocseti63724.2025.11019710
- OpenAlex
- W4410986734
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.