ملف الباحث
Sungmin Cha
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
2025 · arXiv (Cornell University)
Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented -- enabling smaller student models to emulate …