conference-paper Open access

Don’t Rank, Combine! Combining Machine Translation Hypotheses Using Quality Estimation

Research footprint

At a glance

Citations
4
References
0
Comments
0
Paper overview

Abstract

Neural machine translation systems estimate probabilities of target sentences given source sentences, yet these estimates may not align with human preferences.This work introduces QE-fusion, a method that synthesizes translations using a quality estimation metric (QE), which correlates better with human judgments.QE-fusion leverages a pool of candidates sampled from a model and combines spans from different candidates using a QE metric such as COMETKIWI.We compare QE-fusion against beam search and recent reranking techniques, such as Minimum Bayes Risk decoding or QE-reranking.Our method consistently improves translation quality in terms of COMET and BLEURT scores when applied to large language models (LLMs) used for translation (PolyLM, XGLM, Llama2, Mistral, ALMA, and Tower) and to multilingual translation models (NLLB), over five language pairs.Notably, QE-fusion exhibits larger improvements for LLMs due to their ability to generate diverse outputs.We demonstrate that our approach generates novel translations in over half of the cases and consistently outperforms other methods across varying numbers of candidates .Furthermore, we empirically show that QE-fusion scales linearly with the number of candidates in the pool.

Record transparency

Publication details

DOI
10.18653/v1/2024.acl-long.653
OpenAlex
W4402684004
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.