conference-paper

A New Approach of Categorizing Text Documents through Consensus Clustering

Research footprint

At a glance

Citations
1
References
33
Comments
0
Paper overview

Öz

In this paper, we present a new text categorization model using consensus clustering to categorize text documents. Initially, text documents are pre-processed and represented in the form of Term Document Matrix (TDM). Further, consensus clustering is used for text categorization. The consensus clustering has two units: cluster generation and consensus function. In cluster generation unit, we generate clustering results by applying base clustering (K-Means, Fuzzy C-means (FCM) and Intuitionistic Fuzzy C-means (IFCM)) methods. In the consensus function unit, the generated results of base clustering methods are aggregated. Further, voting technique is applied on aggregated results to obtain consolidated cluster. In order to evaluate the effectiveness of the proposed model, experiments are conducted on balanced (20-Newgroups) and unbalanced (Reuters-21578) standard benchmark datasets. We used accuracy, precision, recall and F-measure to assess the performance. The performance of the proposed model is investigated against base clustering methods (K-Means, FCM and IFCM). The experimental result reveals that the proposed model performance is better than base clustering methods. Moreover, consensus clustering eliminates the limitation of base clustering methods and also quality of the final categorization result can be improved.

Record transparency

Publication details

DOI
10.1145/3233347.3233371
OpenAlex
W2891859885
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.