article Open access

TOPNMF: Topic based Document Clustering using Non-negative Matrix Factorization

  • Indian Journal of Science and Technology
  • Indian Society for Education and Environment
Research footprint

At a glance

Citations
2
References
14
Comments
0
Paper overview

Abstract

Objectives: This work focuses on creating targeted content-specific topicbased clusters. They can help users to discover the topics in a set of documents information more efficiently. Methods/Statistical analysis: The Non-negative Matrix Factorization (NMF) based models learn topics by directly decomposing the term-document matrix, which is a bag-of-word matrix representation of a text corpus, into two low-rank factor matrices namely Word-Topic feature Matrix(WTOM) and Document-Topic feature Matrix(DTOM). Topic clusters and Document clusters are extracted from obtained features matrices. This method does not require any statistical distribution and probability. Experiments were carried out on a subset of BBC sport Corpus. Findings: The experimental results indicate that the accuracy of TONMF clusters was observed as 100 percent. Novelty/Applications: NMF often fails to improve the given clustering result as the number of parameters increases linearly with the size of the corpus. The computational complexity of the TOPNMF is better than exact decomposition like Singular Value Decomposition (SVD). Keywords: Topic cluster; Document cluster; Non-negative matrix factorization; K-means clustering; Word cloud

Record transparency

Publication details

DOI
10.17485/ijst/v14i31.1293
OpenAlex
W3202864651
Document type
article
Language
EN
Source
Indian Journal of Science and Technology
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.