Weihao Yu
4 papers in the PaperMetrix corpus
Papers by this author
-
LV-BERT: Exploiting Layer Variety for BERT
2021 · arXiv (Cornell University)
Modern pre-trained language models are mostly built upon backbones stacking self-attention and feed-forward layers in an interleaved order. In this paper, beyond this stereotyped layer pattern, we aim to improve pre-trained models by exploiting layer …
-
Mugs: A Multi-Granular Self-Supervised Learning Framework
2022 · arXiv (Cornell University)
In self-supervised learning, multi-granular features are heavily desired though rarely investigated, as different downstream tasks (e.g., general and fine-grained classification) often require different or multi-granular features, e.g.~fine- or coarse-grained one or their mixture. In this …
-
ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
2020 · arXiv (Cornell University)
Recent powerful pre-trained language models have achieved remarkable performance on most of the popular datasets for reading comprehension. It is time to introduce more challenging datasets to push the development of this field towards more …
-
ConvBERT: Improving BERT with Span-based Dynamic Convolution
2020 · arXiv (Cornell University)
Pre-trained language models like BERT and its variants have recently achieved impressive performance in various natural language understanding tasks. However, BERT heavily relies on the global self-attention block and thus suffers large memory footprint and …