Improving Government-Data Learning via Distributed Clustering Analysis
At a glance
- Citations
- 0
- References
- 4
- Comments
- 0
Abstract
Clustering analysis is a study which is of great value, and the large-scale government-data needed to be handled by cluster analysis is growing increasingly. Efficient analysis techniques of large-scale data need to be adopted to handle the large-scale data. Traditional model of serial programming has serious scalability shortage, which don't satisfy the need of the large-scale government-data handling for computing and storage resources. Distributed computing technology represented by the MapReduce has good scalability, and can greatly improve the execution efficiency of data-intensive algorithm, and give play to the computing power of compute cluster based on general hardware. Based on the background of "data platform for public petition", it aims to study how to combine the cluster analysis technology with the current massive government-data, extracting useful information from the mass characteristics hidden in the data through the cluster analysis technology, which can provide comprehensive analyse for system managers and decision makers. This paper focus on the study of combining basic distributed clustering algorithm and TF-IDF algorithm, developing the cases feature analysis module based on distributed clustering algorithm. Based on distributed clustering algorithm, according to the information of the cases, do clustering analysis of cases according to its characteristics, and then get several hidden information through serveral decisional result.
Publication details
- DOI
- 10.1109/ccbd.2016.053
- OpenAlex
- W2735806961
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.