conference-paper Open access

Data Augmentation for Biomedical Factoid Question Answering

Research footprint

At a glance

Citations
13
References
81
Comments
0
Paper overview

Öz

We study the effect of seven data augmentation (da) methods in factoid question answering, focusing on the biomedical domain, where obtaining training instances is particularly difficult. We experiment with data from the bioasq challenge, which we augment with training instances obtained from an artificial biomedical machine reading comprehension dataset, or via back-translation, information retrieval, word substitution based on word2vec embeddings or masked language modeling, question generation, or extending the given passage with additional context. We show that da can lead to very significant performance gains, even when using large pretrained Transformers, contributing to a broader discussion of if/when da benefits large pretrained models. One of the simplest da methods, word2vec-based word substitution, performed best and is recommended. We release our artificial training instances and code.

Record transparency

Publication details

DOI
10.18653/v1/2022.bionlp-1.6
OpenAlex
W4285214521
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.