article Open access

On learning denoisable student logits

  • Pattern Recognition
  • Elsevier BV
Research footprint

At a glance

Citations
0
References
15
Comments
0
Paper overview

Abstract

Knowledge Distillation (KD) aims to train a student model to mimic the behavior of a more powerful teacher model. In this paper, we reveal that through the lens of diffusion processes, student logits can be statistically treated as a noisy version of teacher logits, and KD helps reduce the noise level of student logits. This insight motivates us to design a framework leveraging KD to produce denoisable student logits that can be further recovered towards teacher logits via a reverse diffusion process. A key advantage of this approach is that the inference-diffusion process can occur in two physical locations and on separate devices, enabling a two-step and distributed inference process. The experimental results show that the derived denoisable student logits achieve comparable or even superior performance to standard KD’s, and the reverse diffusion process achieves a substantial improvement in accuracy, without needing the original image, thus preserving the privacy and security of the original data. Additionally, the logits can be further compressed before transmission, reducing the required bandwidth while achieving comparable overall performance.

Record transparency

Publication details

DOI
10.1016/j.patcog.2026.113684
OpenAlex
W7153292939
Document type
article
Language
EN
Source
Pattern Recognition
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.