article Open access

Synthetic Information Toward Maximum Posterior Ratio for Deep Learning on Imbalanced Data

  • IEEE Transactions on Artificial Intelligence
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
4
References
46
Comments
0
Paper overview

Öz

This work explores how class-imbalanced data affects deep learning and proposes a data balancing technique for mitigation by generating more synthetic data for the minority class. In contrast to random-based oversampling techniques, our approach prioritizes balancing the most informative region by finding high entropy samples. This approach is opportunistic and challenging because well-placed synthetic data points can boost machine learning algorithms’ accuracy and efficiency, whereas poorly-placed ones can cause a higher misclassification rate. In this study, we present an algorithm for maximizing the probability of generating a synthetic sample in the correct region of its class by placing it toward maximizing the class posterior ratio. In addition, to preserve data topology, synthetic data are closely generated within each minority sample neighborhood. Overall, experimental results on forty-one datasets show that our technique significantly outperforms experimental methods in terms of boosting deep-learning performance.

Record transparency

Publication details

DOI
10.1109/tai.2023.3330949
OpenAlex
W4388505581
Document type
article
Language
EN
Source
IEEE Transactions on Artificial Intelligence
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.