conference-paper Open access

Sense-Aware Statistical Machine Translation using Adaptive Context-Dependent Clustering

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Statistical machine translation (SMT) systems use local cues from n-gram translation and language models to select the translation of each source word. Such systems do not explicitly perform word sense disambiguation (WSD), although this would enable them to select translations depending on the hypothesized sense of each word. Previous attempts to constrain word translations based on the results of generic WSD systems have suffered from their limited accuracy. We demonstrate that WSD systems can be adapted to help SMT, thanks to three key achievements: (1)~we consider a larger context for WSD than SMT can afford to consider; (2)~we adapt the number of senses per word to the ones observed in the training data using clustering-based WSD with K-means; and (3)~we initialize sense-clustering with definitions or examples extracted from WordNet. Our WSD system is competitive, and in combination with a factored SMT system improves noun and verb translation from English to Chinese, Dutch, French, German, and Spanish.

Record transparency

Publication details

DOI
10.5281/zenodo.834305
OpenAlex
W4298387127
Document type
conference-paper
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.