preprint
Open access
A Comprehensive Comparison of Pre-training Language Models
Research footprint
At a glance
- Citations
- 0
- References
- 10
- Comments
- 0
Paper overview
Abstract
<p>Recently, the development of pre-trained language models has brought natural language processing (NLP) tasks to the new state-of-the-art. In this paper we explore the efficiency of various pre-trained language models. We pre-train a list of transformer-based models with the same amount of text and the same training steps. The experimental results shows that the most improvement upon the origin BERT is adding the RNN-layer to capture more contextual information for short text understanding. But the conclusion is: There are no remarkable improvement for short text understanding for similar BERT structures. Data-centric method[12] can achieve better performance.</p>
Record transparency
Publication details
- DOI
- 10.36227/techrxiv.14820348
- OpenAlex
- W4252934237
- Document type
- preprint
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.