ملف الباحث
Hongbo Zhang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Co-training and Co-distillation for Quality Improvement and Compression of Language Models
2023 · arXiv (Cornell University)
Knowledge Distillation (KD) compresses computationally expensive pre-trained language models (PLMs) by transferring their knowledge to smaller models, allowing their use in resource-constrained or real-time settings. However, most smaller models fail to surpass the performance of …