Comparative Performance Analysis of Synthetic Minority Oversampling Techniques (SMOTE) on Medical Datasets Based on Extreme Gradient Boosting Estimator
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Medical datasets frequently face issues with class imbalance and redundant features, which can undermine the accuracy and reliability of predictive diagnostic models. This research conducts a performance comparison of four variants of the Synthetic Minority Oversampling Technique (SMOTE)—SMOTE-ENN, Borderline SMOTE, ADASYN, and SMOTE-Tomek Links—when paired with feature selection and the Extreme Gradient Boosting (XGBoost) classifier for predicting breast cancer and heart disease. The publicly available datasets underwent preprocessing to eliminate noise, balance class distributions, and identify the most significant diagnostic features. Model performance was assessed using standard metrics, including accuracy, precision, recall, F1-score, and Cohen's kappa. For the heart disease dataset, the SMOTE-ENN technique produced the highest results, achieving an accuracy of 56.15%, a recall of 44.50%, and an F1-score of 27.46%, which underscored improved detection of minority class cases. Conversely, in the breast cancer dataset, ADASYN, Borderline SMOTE, and SMOTE-Tomek Links showed better performance, reaching an accuracy of 96.49%, an F1-score of 97.30%, a recall of 100%, and a kappa score of 0.9231.
Publication details
- DOI
- 10.5281/zenodo.21353004
- OpenAlex
- W7168249039
- Document type
- article
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.