Improving Bees-based Imputation using Nearest Neighbor for Heuristic Function in Imputing Data
At a glance
- الاستشهادات
- 4
- المراجع
- 10
- Comments
- 0
Abstract
Data imputation is a necessary task to solve missing value problem for better data mining result. The current data imputation with Bees algorithm contains several random procedures including instance selection and feature selection, and the randomness causes inconsistency and swinging result in iteration. Thus, this work proposes to solve them by applying a heuristic function in those procedures from using importance score in selecting attribute to handle and probability in selecting correlated value. These calculations provide the bees with a guidance direction; thus, there are less random processes and should lower inconsistent and swinging results from randomness. From evaluation, the proposed Bees-based imputation obtained higher accuracy than the previous Bees-based and Genetic algorithm-based imputation method from all data sets for all missing data percentage between 10% to 50%. The best improvement in accuracy for 23% in average was found in SPECT data set which consists of only binary type values. For the data sets with values mixing of binary and category type, the proposed method gained about 3-7% improvement in average.
Publication details
- DOI
- 10.1145/3375959.3375974
- OpenAlex
- W3006972075
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.