Semantic Document Clustering Using NLP
At a glance
- Citations
- 0
- References
- 8
- Comments
- 0
Öz
This project explores a semantic-based document clustering system designed to group documents based on the similarity of their content. Unlike traditional keyword-based methods, which rely solely on word frequency, this system leverages Natural Language Processing (NLP) to understand and compare the semantic meaning within documents. Using pre-trained language models such as BERT and Sentence-BERT, each document is converted into a dense vector representation that captures its underlying meaning. These vectors enable precise comparison of documents’ semantic content, allowing for more accurate clustering. The project employs clustering algorithms such as K-Means and DBSCAN, which group documents into clusters based on similarity. Cosine similarity further ensures that related documents are accurately clustered together. Experimental results demonstrate that this approach produces more coherent and contextually relevant clusters compared to traditional techniques, making it an effective solution for applications in content organization, topic analysis, and information retrieval.
Publication details
- DOI
- 10.38124/ijisrt/25may1946
- OpenAlex
- W4411063020
- Document type
- article
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.