conference-paper وصول مفتوح

Towards a Better Metric for Evaluating Question Generation Systems

Research footprint

At a glance

الاستشهادات
99
المراجع
48
Comments
0
Paper overview

Abstract

There has always been criticism for using ngram based similarity metrics, such as BLEU, NIST, etc, for evaluating the performance of NLG systems. However, these metrics continue to remain popular and are recently being used for evaluating the performance of systems which automatically generate questions from documents, knowledge graphs, images, etc. Given the rising interest in such automatic question generation (AQG) systems, it is important to objectively examine whether these metrics are suitable for this task. In particular, it is important to verify whether such metrics used for evaluating AQG systems focus on answerability of the generated question by preferring questions which contain all relevant information such as question type (Wh-types), entities, relations, etc. In this work, we show that current automatic evaluation metrics based on n-gram similarity do not always correlate well with human judgments about answerability of a question. To alleviate this problem and as a first step towards better evaluation metrics for AQG, we introduce a scoring function to capture answerability and show that when this scoring function is integrated with existing metrics, they correlate significantly better with human judgments. The scripts and data developed as a part of this work are made publicly available. 1 Document: In 1648 before the term "genocide" had been coined , the Peace of Westphalia was established to protect ethnic, racial and in some instances religious groups. Possible Question: In which year was the Peace of Westphalia established ?

Record transparency

Publication details

DOI
10.18653/v1/d18-1429
OpenAlex
W2888812214
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.