conference-paper
Open access
Best practices for the human evaluation of automatically generated text
Research footprint
At a glance
- Citations
- 200
- References
- 104
- Comments
- 0
Paper overview
Öz
Currently, there is little agreement as to how Natural Language Generation (NLG) systems should be evaluated, with a particularly high degree of variation in the way that human evaluation is carried out. This paper provides an overview of how human evaluation is currently conducted, and presents a set of best practices, grounded in the literature. With this paper, we hope to contribute to the quality and consistency of human evaluations in NLG.
Record transparency
Publication details
- DOI
- 10.18653/v1/w19-8643
- OpenAlex
- W2992347006
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.