SSKD: Stepwise Self-Knowledge Distillation for Binary Neural Networks in Keyword Spotting
At a glance
- Citations
- 0
- References
- 39
- Comments
- 0
Abstract
The hardware power-aware keyword spotting (KWS) implementation requires small memory footprint, low-complex computation, and high accuracy performances. Binary neural networks (BNNs) naturally satisfy these constraints. They quantize both weights and activations to 1-bit. This reduces storage and replaces most multiply–accumulate operations with bitwise operations. However, such extreme quantization incurs substantial information loss and leaves a noticeable accuracy gap relative to full-precision models. Optimization is also more difficult because the sign function is non-differentiable, and surrogate-gradient updates introduce gradient mismatch. To preserve the hardware benefits of BNNs while alleviating the accuracy degeneration induced by 1-bit quantization, this article addresses the problem from two complementary aspects: Firstly, a Stepwise Self-Knowledge Distillation (SSKD) training approach is proposed to effectively improve the student BNN’s accuracy performance. The SSKD training framework achieves effective supervision for student BNNs. A Stepwise Training Strategy is proposed to optimize the training stability and accuracy. Weight Scaling Factor improves the student’s representational capability. Secondly, an extremely lightweight Binary Temporal Convolutional ResNet (BTC-ResNet) is also proposed. The parameters and calculations inside the network are greatly reduced for the inference. Experiments on the GSCD v1 and GSCD v2 benchmarks demonstrate the effectiveness of our methods for low-power keyword spotting. For the 12-class task, BTC-ResNet14 achieves 97.23% accuracy on GSCD v1 and 97.31% on GSCD v2 with 0.75 Mb parameters and 1.35 M FLOPs. For the 35-class task on GSCD v2, it reaches 95.56% accuracy with 0.76 Mb parameters and 1.35 M FLOPs. These results indicate that our method achieves a competitive accuracy–efficiency balance relative to recent distillation-based BNN KWS baselines reported in the comparative experiments. All these studies are helpful and promising for future KWS deployment on low-power hardware devices.
Publication details
- DOI
- 10.3390/app16042021
- OpenAlex
- W7130352166
- Document type
- article
- Language
- EN
- Source
- Applied Sciences
- Last metadata update
Comments
Log in to join the discussion.