Dealing with Class Imbalance in Machine Learning: Performance Metrics and Data Balancing Methods
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
This work addressed the class imbalance problem in machine learning and identified performance metrics suited to unbalanced datasets. Extensive experiments were conducted on both simulated and real-life data using various machine learning algorithms and data balancing techniques. The findings highlighted the limitations of accuracy and emphasized that while ROC-AUC was a more suitable metric for evaluating model performance, balancing the dataset was essential for obtaining reliable results. Furthermore, the stability of the Matthews Correlation Coefficient (MCC) before and after data balancing underscored its robustness as a performance measure. Additionally, the study revealed that while various data-balancing techniques had a similar impact on improving machine learning model performance, undersampling consistently outperformed other methods by enhancing the performance of the Random Forest model in terms of the ROC-AUC metric across all cases.
Publication details
- DOI
- 10.22541/au.175666516.63936610/v1
- OpenAlex
- W4413862839
- Document type
- preprint
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.