conference-paper

Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech Representation

Research footprint

At a glance

Citations
4
References
34
Comments
0
Paper overview

Abstract

Predicting the altered acoustic frames is an effective way of self-supervised learning for speech representation.However, it is challenging to prevent the pretrained model from overfitting.In this paper, we proposed to introduce two dropout regularization methods into the pretraining of transformer encoder: (1) attention dropout, (2) layer dropout.Both of the two dropout methods encourage the model to utilize global speech information, and avoid just copying local spectrum features when reconstructing the masked frames.We evaluated the proposed methods on phoneme classification and speaker recognition tasks.The experiments demonstrate that our dropout approaches achieve competitive results, and improve the performance of classification accuracy on downstream tasks.

Record transparency

Publication details

DOI
10.21437/interspeech.2021-1066
OpenAlex
W3196798358
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.