preprint
Open access
The Effects of In-domain Corpus Size on pre-training BERT
Research footprint
At a glance
- Citations
- 2
- References
- 0
- Comments
- 0
Paper overview
Abstract
Many prior language modeling efforts have shown that pre-training on an in-domain corpus can significantly improve performance on downstream domain-specific NLP tasks. However, the difficulties associated with collecting enough in-domain data might discourage researchers from approaching this pre-training task. In this paper, we conducted a series of experiments by pre-training Bidirectional Encoder Representations from Transformers (BERT) with different sizes of biomedical corpora. The results demonstrate that pre-training on a relatively small amount of in-domain data (4GB) with limited training steps, can lead to better performance on downstream domain-specific NLP tasks compared with fine-tuning models pre-trained on general corpora.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.2212.07914
- OpenAlex
- W4311728218
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.