On learning denoisable student logits
At a glance
- Citations
- 0
- References
- 15
- Comments
- 0
Abstract
Knowledge Distillation (KD) aims to train a student model to mimic the behavior of a more powerful teacher model. In this paper, we reveal that through the lens of diffusion processes, student logits can be statistically treated as a noisy version of teacher logits, and KD helps reduce the noise level of student logits. This insight motivates us to design a framework leveraging KD to produce denoisable student logits that can be further recovered towards teacher logits via a reverse diffusion process. A key advantage of this approach is that the inference-diffusion process can occur in two physical locations and on separate devices, enabling a two-step and distributed inference process. The experimental results show that the derived denoisable student logits achieve comparable or even superior performance to standard KD’s, and the reverse diffusion process achieves a substantial improvement in accuracy, without needing the original image, thus preserving the privacy and security of the original data. Additionally, the logits can be further compressed before transmission, reducing the required bandwidth while achieving comparable overall performance.
Publication details
- DOI
- 10.1016/j.patcog.2026.113684
- OpenAlex
- W7153292939
- Document type
- article
- Language
- EN
- Source
- Pattern Recognition
- Last metadata update
Comments
Log in to join the discussion.