article

A Knowledge Distillation-Based Approach to Speech Emotion Recognition

  • IEEE Transactions on Affective Computing
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
3
References
45
Comments
0
Paper overview

Abstract

Due to rapid advancements in deep learning, Transformer-based architectures have proven effective in speech emotion recognition (SER), largely due to their ability to model long-term dependencies more effectively than recurrent networks. The current Transformer architecture is not well-suited for SER because its large parameter number demands significant computational resources, making it less feasible in environments with limited resources. Furthermore, its application to SER is limited because human emotions, which are expressed in long segments of continuous speech, are inherently complex and ambiguous. Therefore, designing specialized Transformer models tailored for SER is essential. To address these challenges, we propose a novel knowledge distillation framework that combines meta-knowledge and curriculum-based distillation. Specifically, we fine-tune the teacher model to optimize it for the SER task. For the student model, we embed individual sequence time points into variable tokens, which are used to aggregate the global speech representation. Additionally, we combine supervised contrastive and cross-entropy loss to increase the inter-class distance between learnable features. Finally, we optimize the student model using both meta-knowledge and the curriculum-based distillation framework. Experimental results on two benchmark datasets, IEMOCAP and MELD, demonstrate that our method performs competitively with state-of-the-art approaches in SER.

Record transparency

Publication details

DOI
10.1109/taffc.2025.3574178
OpenAlex
W4411019804
Document type
article
Language
EN
Source
IEEE Transactions on Affective Computing
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.