article وصول مفتوح

The Effect of Preprocessing on Short Document Clustering

  • Repository KITopen (Karlsruhe Institute of Technology)
  • Karlsruhe Institute of Technology
Research footprint

At a glance

الاستشهادات
5
المراجع
0
Comments
0
Paper overview

Abstract

Natural Language Processing has become a common tool to extract relevant information from unstructured data. Messages in social media, customer reviews, and military messages are all very short and therefore harder to handle than longer texts. Document clustering is essential in gaining insight from these unlabeled texts and is typically performed after some preprocessing steps. Preprocessing often removes words. This can become risky in short texts, where the main message is made of only a few words. The effect of preprocessing and feature extraction on these short documents is therefore analyzed in this paper. Six different levels of text normalization are combined with four different feature extraction methods. These setting are all applied on K-means clustering and tested on three different datasets. Anticipated results can not be concluded, however other findings are insightful in terms of the connection between text cleaning and feature extraction.

Record transparency

Publication details

DOI
10.5445/ksp/1000098011/01
OpenAlex
W3043786142
Document type
article
Language
EN
Source
Repository KITopen (Karlsruhe Institute of Technology)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.