conference-paper

The Effects of Random Undersampling for Big Data Medicare Fraud Detection

Research footprint

At a glance

الاستشهادات
20
المراجع
26
Comments
0
Paper overview

Abstract

We show it is possible to obtain better classification performance for experiments involving highly imbalanced Big Data with the application of data sampling techniques. We apply Random Undersampling to a publicly available Medicare insurance claims dataset of about 175 million records. We find that Random Undersampling significantly improves the classification performance of Extremely Randomized Trees and XGBoost learners in a Medicare Big Data fraud classification task. We employ Random Undersampling to evaluate performance at multiple minority:majority class ratios. According to the outcome of a Tukey’s Honestly Significant Difference test, we find Random Undersampling to the 1:9 or 1:27 class ratios yields the best performance, providing Area Under the Receiver Operating Characteristic scores of over 0.97. Models built with undersampled Big Data require significantly less time to train. Our contribution is to prove the effectiveness of Random Undersampling in classifying Medicare Big Data. Our review of related work shows we are the first to apply Random Undersampling to data on this scale in order to prove one can obtain better performance. To the best of our knowledge, we are the first to perform experiments with the latest Medicare Part D data, made available in 2021.

Record transparency

Publication details

DOI
10.1109/sose55356.2022.00023
OpenAlex
W4312674672
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.