conference-paper

An Empirical Comparison of Three Ensemble Methods for Medical Data Mining with Apache Spark

Research footprint

At a glance

Citations
2
References
26
Comments
0
Paper overview

Öz

Medical data in various organizational forms are voluminous and heterogeneous, so it is highly meaningful to utilize parallel computing platforms to speed up the data mining procedure. In addition, ensemble methods which combine different weak classifiers together can improve classification accuracy on parallel computing platforms. In this paper, three ensemble methods (Bagging, AdaBoost and Logit Boost) with logistic regression as the weak classifier are implemented with Apache Spark for achieving better parallel computing performance and taking full advantage of RDD. And a series of experiments are carried out in different execution modes to evaluate and compare the classification performance and the parallelism of these ensemble methods. Experimental results indicate that although Bagging is slightly inferior to AdaBoost and Logit Boost in classification accuracy, it achieves better parallelism than the other two methods. Finally, selection criteria of these ensemble methods are presented in accordance with specific medical application scenario.

Record transparency

Publication details

DOI
10.1109/uic-atc-scalcom-cbdcom-iop.2015.175
OpenAlex
W2478066838
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.