conference-paper Open access

ILDAE: Instance-Level Difficulty Analysis of Evaluation Data

  • Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Research footprint

At a glance

Citations
17
References
49
Comments
0
Paper overview

Abstract

Knowledge of difficulty level of questions helps a teacher in several ways, such as estimating students' potential quickly by asking carefully selected questions and improving quality of examination by modifying trivial and hard questions. Can we extract such benefits of instance difficulty in Natural Language Processing? To this end, we conduct Instance-Level Difficulty Analysis of Evaluation data (ILDAE) in a largescale setup of 23 datasets and demonstrate its five novel applications: 1) conducting efficientyet-accurate evaluations with fewer instances saving computational cost and time, 2) improving quality of existing evaluation datasets by repairing erroneous and trivial instances, 3) selecting the best model based on application requirements, 4) analyzing dataset characteristics for guiding future data creation, 5) estimating Out-of-Domain performance reliably. Comprehensive experiments for these applications lead to several interesting results, such as evaluation using just 5% instances (selected via ILDAE) achieves as high as 0.93 Kendall correlation with evaluation using complete dataset and computing weighted accuracy using difficulty scores leads to 5.2% higher correlation with Out-of-Domain performance. We release the difficulty scores 1 and hope our work will encourage research in this important yet understudied field of leveraging instance difficulty in evaluations.

Record transparency

Publication details

DOI
10.18653/v1/2022.acl-long.240
OpenAlex
W4285111874
Document type
conference-paper
Language
EN
Source
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.