article Open access

Refining Synthesized Speech Using Speaker Information and Phone Masking for Data Augmentation of Speech Recognition

  • IEEE/ACM Transactions on Audio Speech and Language Processing
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
3
References
57
Comments
0
Paper overview

Abstract

While end-to-end automatic speech recognition (ASR) has shown impressive performance, it requires a huge amount of speech and transcription data. The conversion of domain-matched text to speech (TTS) has been investigated as one approach to data augmentation. The quality and diversity of the synthesized speech are critical in this approach. To ensure quality, a neural vocoder is widely used to generate speech waveforms in conventional studies, but it requires a huge amount of computation and another conversion to spectral-domain features such as the log-Mel filterbank (lmfb) output typically used for ASR. In this study, we explore the direct refinement of these features. Unlike conventional speech enhancement, we can use information on the ground-truth phone sequences of the speech and designated speaker to improve the quality and diversity. This process is realized as a Mel-to-Mel network, which can be placed after a text-to-Mel synthesis system such as FastSpeech 2. These two networks can be trained jointly. Moreover, semantic masking is applied to the lmfb features for robust training. Experimental evaluations demonstrate the effect of phone information, speaker information, and semantic masking. For speaker information, x-vector performs better than the simple speaker embedding. The proposed method achieves even better ASR performance with a much shorter computation time than the conventional method using a vocoder.

Record transparency

Publication details

DOI
10.1109/taslp.2024.3451982
OpenAlex
W4402187093
Document type
article
Language
EN
Source
IEEE/ACM Transactions on Audio Speech and Language Processing
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.