conference-paper

Optimizing Speech Recognition Models with Effective Confidence Calibrations

Research footprint

At a glance

Citations
0
References
20
Comments
0
Paper overview

Abstract

In the domain of speech recognition, deep learning models have achieved tremendous success. Despite the increasing predictive performance of these models, there has been limited focus on uncertainty estimation and calibration, leading to unreliable confidence. To address this challenge, we employed Sparse Gaussian Process Attention (SGPA), which performs Bayesian inference directly within the output space of the Multi-Head Attention (MHA) blocks in Transformers to calibrate its uncertainty. This method replaces the scaled dot-product operation with an effective symmetric kernel and utilizes Sparse Gaussian Process (SGP) techniques to approximate the posterior distribution of the MHA outputs. Experimental results in speech recognition tasks demonstrate that models incorporating SGPA not only maintain their predictive accuracy but also significantly enhance the reliability of their output confidence.

Record transparency

Publication details

DOI
10.1109/cac63892.2024.10865370
OpenAlex
W4407451547
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.