article

Multimodular Text Normalization of Dutch User-Generated Content

  • ACM Transactions on Intelligent Systems and Technology
  • Association for Computing Machinery
Research footprint

At a glance

Citations
37
References
68
Comments
0
Paper overview

Abstract

As social media constitutes a valuable source for data analysis for a wide range of applications, the need for handling such data arises. However, the nonstandard language used on social media poses problems for natural language processing (NLP) tools, as these are typically trained on standard language material. We propose a text normalization approach to tackle this problem. More specifically, we investigate the usefulness of a multimodular approach to account for the diversity of normalization issues encountered in user-generated content (UGC). We consider three different types of UGC written in Dutch (SNS, SMS, and tweets) and provide a detailed analysis of the performance of the different modules and the overall system. We also apply an extrinsic evaluation by evaluating the performance of a part-of-speech tagger, lemmatizer, and named-entity recognizer before and after normalization.

Record transparency

Publication details

DOI
10.1145/2850422
OpenAlex
W2371227879
Document type
article
Language
EN
Source
ACM Transactions on Intelligent Systems and Technology
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.