Optimization and realization of parallel frequent item set mining algorithm
At a glance
- Citations
- 0
- References
- 12
- Comments
- 0
Abstract
Associative data mining is the research hotspot in the field of big data, and frequent item sets mining is an important step in the analysis of associative data. This paper focuses on analyzing the frequent item sets mining algorithm based on Apriori parallel algorithm. The paper has found two shortages of Apriori parallel algorithm: one is that the key value pair are too many, another is that in the combiner stage, it occupies two much memory. Therefore, we propose an optimized algorithm. In the optimization algorithm, candidate item sets and local count information are saved in memory, greatly reducing the number of generated keys. Meanwhile, in the short length frequent item sets mining, the method of reducing the number of scanning transaction data without generating candidate item sets can improve the algorithm efficiency. We do the experiments in the Hadoop platform to testify the performance of the proposed optimized algorithm. The experiments demonstrate that the time and I/O of the optimized algorithm have been improved greatly, compared with the non-optimized algorithm.
Publication details
- DOI
- 10.1109/icalip.2016.7846585
- OpenAlex
- W2586472337
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.