Kaitao Song
4 papers in the PaperMetrix corpus
Papers by this author
-
Task-Agnostic and Adaptive-Size BERT Compression
2021
While pre-trained language models such as BERT and RoBERTa have achieved impressive results on various natural language processing tasks, they have huge numbers of parameters and suffer from huge computational and memory costs, which make …
-
DiffusionNER: Boundary Diffusion for Named Entity Recognition
2023 · arXiv (Cornell University)
In this paper, we propose DiffusionNER, which formulates the named entity recognition task as a boundary-denoising diffusion process and thus generates named entities from noisy spans. During training, DiffusionNER gradually adds noises to the golden …
-
MASS: Masked Sequence to Sequence Pre-training for Language Generation
2019 · arXiv (Cornell University)
Pre-training and fine-tuning, e.g., BERT, have achieved great success in language understanding by transferring knowledge from rich-resource pre-training task to the low/zero-resource downstream tasks. Inspired by the success of BERT, we propose MAsked Sequence to …
-
MPNet: Masked and Permuted Pre-training for Language Understanding
2020 · arXiv (Cornell University)
BERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this …