ملف الباحث
Gongfan Fang
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Adversarial Self-Supervised Data-Free Distillation for Text Classification
2020 · arXiv (Cornell University)
Large pre-trained transformer-based language models have achieved impressive results on a wide range of NLP tasks. In the past few years, Knowledge Distillation(KD) has become a popular paradigm to compress a computationally expensive model to …
-
Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data
2021 · arXiv (Cornell University)
Knowledge distillation~(KD) aims to craft a compact student model that imitates the behavior of a pre-trained teacher in a target domain. Prior KD approaches, despite their gratifying results, have largely relied on the premise that …