conference-paper

An Ensemble-Based Framework for Enhanced Missing Data Imputation

Research footprint

At a glance

Citations
3
References
16
Comments
0
Paper overview

Abstract

Missing data is a pervasive issue in many real-world datasets, often leading to biased estimates, reduced statistical power, and invalid conclusions if not properly addressed. Traditional imputation methods, such as mean, median, and K-nearest neighbors (KNN), each have their advantages but often fail to provide consistently accurate estimates across different datasets and missing data patterns. Advanced methods, including IterativeImputer and matrix factorization-based SoftImpute, offer improvements but still face limitations. In this paper, we propose a novel ensemble-based imputation method that synergistically combines the strengths of KNN, IterativeImputer, mean, median, and SoftImpute. The ensemble method leverages the performance of each individual technique to provide the most accurate imputation for each missing value. We conducted extensive evaluations on a dataset with artificially introduced 2%, 5%, 10% and 20% missing data. Our results, analysed using root mean square error (RMSE), demonstrate that the proposed ensemble method consistently outperforms each individual imputation method across various levels of missing data. This research highlights the effectiveness of an ensemble approach in enhancing data quality and reliability, offering a robust solution for practitioners dealing with incomplete datasets.

Record transparency

Publication details

DOI
10.1109/csnt64827.2025.10967640
OpenAlex
W4409724906
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.