Data Augmentation for Biomedical Factoid Question Answering
At a glance
- Citations
- 13
- References
- 81
- Comments
- 0
Öz
We study the effect of seven data augmentation (da) methods in factoid question answering, focusing on the biomedical domain, where obtaining training instances is particularly difficult. We experiment with data from the bioasq challenge, which we augment with training instances obtained from an artificial biomedical machine reading comprehension dataset, or via back-translation, information retrieval, word substitution based on word2vec embeddings or masked language modeling, question generation, or extending the given passage with additional context. We show that da can lead to very significant performance gains, even when using large pretrained Transformers, contributing to a broader discussion of if/when da benefits large pretrained models. One of the simplest da methods, word2vec-based word substitution, performed best and is recommended. We release our artificial training instances and code.
Publication details
- DOI
- 10.18653/v1/2022.bionlp-1.6
- OpenAlex
- W4285214521
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.