conference-paper

A Comparison of Several Word Clustering Models

Research footprint

At a glance

Citations
1
References
6
Comments
0
Paper overview

Abstract

Sparse-data problem is a main issue that influences the performances of statistical language models; statistical language model based on word classes is an effective method to solve sparse-data problems. This paper presents a definition of word similarity by utilizing mutual information of adjoining words, and gives the definition of word set similarity based on word similarity, and puts forward a bottom-up hierarchical word clustering algorithm which can get global optimum. Experimental results show that the word clustering algorithm is of high executing speed and have good clustering performances. We then interpolated the class-based models with the word-based models and found that it mitigates remaining sparse-data problems of statistical language models.

Record transparency

Publication details

DOI
10.1109/iaeac47372.2019.8997887
OpenAlex
W3005927046
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.