article Open access

Document Preprocessing with TF-IDF to Improve the Polarity Classification Performance of Unstructured Sentiment Analysis

  • Kinetik Game Technology Information System Computer Network Computing Electronics and Control
  • Muhammadiyah University of Malang
Research footprint

At a glance

Citations
24
References
25
Comments
0
Paper overview

Abstract

Sentiment analysis in terms of polarity classification is very important in everyday life, with the existence of polarity, many people can find out whether the respected document has positive or negative sentiment so that it can help in choosing and making decisions. Sentiment analysis usually done manually. Therefore, an automatic sentiment analysis classification process is needed. However, it is rare to find studies that discuss extraction features and which learning models are suitable for unstructured sentiment analysis types with the Amazon food review case. This research explores some extraction features such as Word Bags, TF-IDF, Word2Vector, as well as a combination of TF-IDF and Word2Vector with several machine learning models such as Random Forest, SVM, KNN and Naïve Bayes to find out a combination of feature extraction and learning models that can help add variety to the analysis of polarity sentiments. By assisting with document preparation such as html tags and punctuation and special characters, using snowball stemming, TF-IDF results obtained with SVM are suitable for obtaining a polarity classification in unstructured sentiment analysis for the case of Amazon food review with a performance result of 87,3 percent.

Record transparency

Publication details

DOI
10.22219/kinetik.v5i3.1066
OpenAlex
W3082312505
Document type
article
Language
EN
Source
Kinetik Game Technology Information System Computer Network Computing Electronics and Control
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.