article وصول مفتوح

Sentiment analysis datasets: a review of trends, challenges, and future opportunities

  • PeerJ Computer Science
  • PeerJ, Inc.
Research footprint

At a glance

الاستشهادات
0
المراجع
145
Comments
0
Paper overview

Abstract

Sentiment Analysis (SA) is a core area within Natural Language Processing (NLP) that focuses on computationally identifying and interpreting subjective information. The performance and generalizability of SA models are closely linked to the characteristics and quality of the datasets used for training and evaluation. Although extensive research has advanced algorithmic techniques in SA, the literature offers comparatively limited consolidated analyses of the datasets that underpin benchmarking and empirical progress. To address this gap, this review provides an updated and comprehensive synthesis focused exclusively on SA datasets. We present a structured taxonomy based on labeling strategies, domain specificity, data sources, granularity levels, and modalities, enabling a clearer understanding of dataset typologies and their implications for SA research. The study analyzes key limitations in existing datasets and discusses emerging directions in dataset development, including multimodality, cross-lingual expansion, and synthetic data augmentation. Furthermore, we outline best practices for dataset annotation, validation, and benchmarking to support the development of more robust and transferable SA systems. By consolidating and examining the diverse landscape of SA datasets, this review aims to serve as a valuable resource for researchers and practitioners, offering guidance for informed dataset selection and laying a foundation for future advancements in SA dataset design and utilization.

Record transparency

Publication details

DOI
10.7717/peerj-cs.3806
OpenAlex
W7163666565
Document type
article
Language
EN
Source
PeerJ Computer Science
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.