conference-paper Open access

A Data Cartography based MixUp for Pre-trained Language Models

  • Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Research footprint

At a glance

Citations
5
References
17
Comments
0
Paper overview

Abstract

MixUp is a data augmentation strategy where additional samples are generated during training by combining random pairs of training samples and their labels. However, selecting random pairs is not potentially an optimal choice. In this work, we propose TD-MixUp, a novel MixUp strategy that leverages Training Dynamics and allows more informative samples to be combined for generating new data samples. Our proposed TD-MixUp first measures confidence, variability, We empirically validate that our method not only achieves competitive performance using a smaller subset of the training data compared with strong baselines, but also yields lower expected calibration error on the pre-trained language model, BERT, on both in-domain and out-of-domain settings in a wide range of NLP tasks. We publicly release our code. 1

Record transparency

Publication details

DOI
10.18653/v1/2022.naacl-main.314
OpenAlex
W4229456195
Document type
conference-paper
Language
EN
Source
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.