An Ensemble-Based Framework for Enhanced Missing Data Imputation
At a glance
- Citations
- 3
- References
- 16
- Comments
- 0
Abstract
Missing data is a pervasive issue in many real-world datasets, often leading to biased estimates, reduced statistical power, and invalid conclusions if not properly addressed. Traditional imputation methods, such as mean, median, and K-nearest neighbors (KNN), each have their advantages but often fail to provide consistently accurate estimates across different datasets and missing data patterns. Advanced methods, including IterativeImputer and matrix factorization-based SoftImpute, offer improvements but still face limitations. In this paper, we propose a novel ensemble-based imputation method that synergistically combines the strengths of KNN, IterativeImputer, mean, median, and SoftImpute. The ensemble method leverages the performance of each individual technique to provide the most accurate imputation for each missing value. We conducted extensive evaluations on a dataset with artificially introduced 2%, 5%, 10% and 20% missing data. Our results, analysed using root mean square error (RMSE), demonstrate that the proposed ensemble method consistently outperforms each individual imputation method across various levels of missing data. This research highlights the effectiveness of an ensemble approach in enhancing data quality and reliability, offering a robust solution for practitioners dealing with incomplete datasets.
Publication details
- DOI
- 10.1109/csnt64827.2025.10967640
- OpenAlex
- W4409724906
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.