article Open access

A Structured Review of the Validity of BLEU

  • Computational Linguistics
  • Association for Computational Linguistics
Research footprint

At a glance

Citations
361
References
18
Comments
0
Paper overview

Abstract

The BLEU metric has been widely used in NLP for over 15 years to evaluate NLP systems, especially in machine translation and natural language generation. I present a structured review of the evidence on whether BLEU is a valid evaluation technique—in other words, whether BLEU scores correlate with real-world utility and user-satisfaction of NLP systems; this review covers 284 correlations reported in 34 papers. Overall, the evidence supports using BLEU for diagnostic evaluation of MT systems (which is what it was originally proposed for), but does not support using BLEU outside of MT, for evaluation of individual texts, or for scientific hypothesis testing.

Record transparency

Publication details

DOI
10.1162/coli_a_00322
OpenAlex
W2806532810
Document type
article
Language
EN
Source
Computational Linguistics
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.