ملف الباحث
Qibin Hou
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
LV-BERT: Exploiting Layer Variety for BERT
2021 · arXiv (Cornell University)
Modern pre-trained language models are mostly built upon backbones stacking self-attention and feed-forward layers in an interleaved order. In this paper, beyond this stereotyped layer pattern, we aim to improve pre-trained models by exploiting layer …