conference-paper Open access

Dice Loss for Data-imbalanced NLP Tasks

Research footprint

At a glance

Citations
595
References
73
Comments
0
Paper overview

Abstract

Many NLP tasks such as tagging and machine reading comprehension (MRC) are faced with the severe data imbalance issue: negative examples significantly outnumber positive ones, and the huge number of easy-negative examples overwhelms training. The most commonly used cross entropy criteria is actually accuracy-oriented, which creates a discrepancy between training and test. At training time, each training instance contributes equally to the objective function, while at test time F1 score concerns more about positive examples.

Record transparency

Publication details

DOI
10.18653/v1/2020.acl-main.45
OpenAlex
W3034328552
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.