Data augmentation for speaker verification
At a glance
- الاستشهادات
- 3
- المراجع
- 23
- Comments
- 0
Abstract
Data augmentation is a hot issue in neural network training. In this paper, we investigate data augmentation for speaker recognition, and we propose two data augmentation methods to enhance the performance of neural network system. One of which is spectral augmentation. Spectral augmentation is a newly proposed data augmentation method which applied to speech recognition and got state-of-the-art performance, by masking blocks of frequency channels(F-mask), and/or by masking blocks of time steps(T-mask). We also investigate the method of speed perturbation, which adjusts the time-scale of a given audio signal without altering its pitch content. Experimental results show that both two methods can boost the performance. By combining TF-mask with speed perturbation, we can obtain more than 5.2% and 8.7% relative improvements over the baseline systems in the Vox-H and WX tasks.
Publication details
- DOI
- 10.1145/3573428.3573649
- OpenAlex
- W4327521002
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.