conference-paper Open access

Q2: : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering

  • Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Research footprint

At a glance

Citations
106
References
64
Comments
0
Paper overview

Abstract

Neural knowledge-grounded generative models for dialogue often produce content that is factually inconsistent with the knowledge they rely on, making them unreliable and limiting their applicability. Inspired by recent work on evaluating factual consistency in abstractive summarization, we propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering. Our metric, denoted Q 2 , compares answer spans using natural language inference (NLI), instead of token-based matching as done in previous work. To foster proper evaluation, we curate a novel dataset of dialogue system outputs for the Wizard-of-Wikipedia dataset, manually annotated for factual consistency. We perform a thorough meta-evaluation of Q 2 against other metrics using this dataset and two others, where it consistently shows higher correlation with human judgements.

Record transparency

Publication details

DOI
10.18653/v1/2021.emnlp-main.619
OpenAlex
W3153947101
Document type
conference-paper
Language
EN
Source
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.