article Open access

Addressing the Binning Problem in Calibration Assessment through Scalar Annotations

  • Transactions of the Association for Computational Linguistics
  • Association for Computational Linguistics
Research footprint

At a glance

Citations
1
References
67
Comments
0
Paper overview

Öz

Abstract Computational linguistics models commonly target the prediction of discrete—categorical—labels. When assessing how well-calibrated these model predictions are, popular evaluation schemes require practitioners to manually determine a binning scheme: grouping labels into bins to approximate true label posterior. The problem is that these metrics are sensitive to binning decisions. We consider two solutions to the binning problem that apply at the stage of data annotation: collecting either distributed (redundant) labels or direct scalar value assignment. In this paper, we show that although both approaches address the binning problem by evaluating instance-level calibration, direct scalar assignment is significantly more cost-effective. We provide theoretical analysis and empirical evidence to support our proposal for dataset creators to adopt scalar annotation protocols to enable a higher-quality assessment of model calibration.

Record transparency

Publication details

DOI
10.1162/tacl_a_00636
OpenAlex
W4391852929
Document type
article
Language
EN
Source
Transactions of the Association for Computational Linguistics
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.