Explainable Multilevel Ensemble for Breast Cancer Prediction with Gene Expression Data
At a glance
- Citations
- 0
- References
- 14
- Comments
- 0
Abstract
Early detection of breast cancer enhances survival and treatment success. Gene expression analysis helps by identifying significant genetic patterns for accurate diagnosis and prognosis. However, the presence of redundant features in high-dimensional gene expression data poses significant challenges for classification, often reducing the performance and interpretability of traditional machine learning models. To address this, we propose an explainable multilevel ensemble framework that effectively combines the strengths of feature selection and ensemble classification. Specifically, SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-Agnostic Explanations) are independently applied to identify the most informative genes based on their contribution to model predictions. The features combinedly used to train an ensemble classifier that integrates XGBoost and Random Forest. This strategy enhances the predictive accuracy, ensures greater transparency and interpretability through Explainable Artificial Intelligence (XAI). Experimental results on publicly available breast cancer gene expression datasets demonstrate that the proposed framework outperforms traditional models in terms of accuracy, sensitivity, specificity, and F1-curve. Furthermore, the use of explainable methods enables deeper biological insights by identifying the most influential genes contributing to cancer classification, highlighting the framework’s potential as a valuable decision-support tool in precision oncology.
Publication details
- DOI
- 10.1109/icaiss61471.2025.11041845
- OpenAlex
- W4411599859
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.