conference-paper Open access

Determining Gains Acquired from Word Embedding Quantitatively Using Discrete Distribution Clustering

Research footprint

At a glance

Citations
14
References
41
Comments
0
Paper overview

Abstract

Word embeddings have become widelyused in document analysis. While a large number of models for mapping words to vector spaces have been developed, it remains undetermined how much net gain can be achieved over traditional approaches based on bag-of-words. In this paper, we propose a new document clustering approach by combining any word embedding with a state-of-the-art algorithm for clustering empirical distributions. By using the Wasserstein distance between distributions, the word-to-word semantic relationship is taken into account in a principled way. The new clustering method is easy to use and consistently outperforms other methods on a variety of data sets.

Record transparency

Publication details

DOI
10.18653/v1/p17-1169
OpenAlex
W2739844977
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.