Optimizing Speech Recognition Models with Effective Confidence Calibrations
At a glance
- Citations
- 0
- References
- 20
- Comments
- 0
Abstract
In the domain of speech recognition, deep learning models have achieved tremendous success. Despite the increasing predictive performance of these models, there has been limited focus on uncertainty estimation and calibration, leading to unreliable confidence. To address this challenge, we employed Sparse Gaussian Process Attention (SGPA), which performs Bayesian inference directly within the output space of the Multi-Head Attention (MHA) blocks in Transformers to calibrate its uncertainty. This method replaces the scaled dot-product operation with an effective symmetric kernel and utilizes Sparse Gaussian Process (SGP) techniques to approximate the posterior distribution of the MHA outputs. Experimental results in speech recognition tasks demonstrate that models incorporating SGPA not only maintain their predictive accuracy but also significantly enhance the reliability of their output confidence.
Publication details
- DOI
- 10.1109/cac63892.2024.10865370
- OpenAlex
- W4407451547
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.