conference-paper

Comparison of Data Augmentation Techniques on Filipino ASR for Children’s Speech

Research footprint

At a glance

Citations
3
References
16
Comments
0
Paper overview

Abstract

While recent advances in automatic speech recognition (ASR) systems involve neural architectures and large acoustic models with latent feature representations, problem arises for low-resource languages- especially when dealing with children’s speech that has very limited data. In this paper, recent data augmentation techniques such as spectral warping (SW), vocal tract length perturbation (VTLP), spectrogram augmentation (SpecAug), and MaskCycleGAN-VC were used on the Tagalog ISIP06 corpus. Moreover, experiments were designed to determine the optimal parameters for each technique and to evaluate the systems using different combinations of DA in terms of the word error rate (WER) and relative improvement (RI) upon testing with actual children’s speech data. Based on the results, the combined augmented data which yielded the best-performing system recorded the lowest WER of 11.72% corresponding to a 43.55% relative improvement with respect to the baseline system.

Record transparency

Publication details

DOI
10.1109/sped59241.2023.10314952
OpenAlex
W4388692821
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.