Determining the Most Accurate Text Classifier Model for Predicting Whether Online Product Reviews are Human or Computer Generated
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
In recent years, fake online product reviews have become an increasingly prevalent issue for e-commerce platforms (McCluskey, 2022). This project attempts to address this issue of pervasive fake reviews negatively impacting e-commerce platforms by focusing on the automated detection of computer-generated reviews. The study evaluates several text classifier algorithms, including Logistic Regression, Support Vector, Decision Tree, K-Nearest Neighbors (KNN), Random Forest, Extra Trees, AdaBoost, Bagging, and Gradient Boosting, and compares their performances in distinguishing computer-generated reviews from genuine reviews. These classifiers are assessed based on their accuracy and F1 scores (the harmonic mean of their precision and recall scores) in distinguishing between human and computer-generated reviews after being trained on a dataset comprising of approximately 20,000 computer-generated and approximately 20,000 human-written reviews, which have been transformed into numerical vectors using TF-IDF vectorization. Results indicate that Logistic Regression consistently outperforms other classifiers, demonstrating robust accuracy and F1 scores across trials. K-Nearest Neighbors classifier shows the poorest performance, likely due to challenges in highdimensional text data. Ensemble methods, such as Random Forest and Extra Trees, deliver notable success, leveraging multiple decision trees to enhance predictive performance. AdaBoost and Gradient Boosting also demonstrate competitive results, showcasing the capabilities of adaptive boosting. The study concludes that Logistic Regression is the most accurate classifier for detecting computer-generated reviews, offering insights into its simplicity and effectiveness in capturing linear patterns within the data. Further exploration could involve tuning hyperparameters for models with poor performance, and exploration with even more models.
Publication details
- DOI
- 10.70121/001c.121720
- OpenAlex
- W4401245795
- Document type
- article
- Language
- EN
- Source
- Scholarly review .
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.