conference-paper Open access

Pulling Out the Stops: Rethinking Stopword Removal for Topic Models

Research footprint

At a glance

Citations
189
References
11
Comments
0
Paper overview

Abstract

It is often assumed that topic models benefit from the use of a manually curated stopword list. Constructing this list is timeconsuming and often subject to user judgments about what kinds of words are important to the model and the application. Although stopword removal clearly affects which word types appear as most probable terms in topics, we argue that this improvement is superficial, and that topic inference benefits little from the practice of removing stopwords beyond very frequent terms. Removing corpus-specific stopwords after model inference is more transparent and produces similar results to removing those words prior to inference.

Record transparency

Publication details

DOI
10.18653/v1/e17-2069
OpenAlex
W2742034229
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.