Investigating the impact of the training data volume for robust speech recognition using multi-task learning
At a glance
- Citations
- 1
- References
- 23
- Comments
- 0
Abstract
Dealing with speech corrupted by noise and reverberation is still an issue for automatic speech recognition. To address this, a solution that can be combined with multi-style learning consists of using multi-task learning, where the acoustic model is trained to solve one main task and at least one auxiliary task simultaneously. In noisy and reverberant environment using clean-speech features estimation as auxiliary task has proven its efficiency, while the main task is speech recognition. Still, recognizing speech in these degraded conditions is all the more difficult when the amount of available data during training is limited. Thus, in this paper, we evaluate robust speech recognition based on multi-task learning when the amount of training data is gradually reduced. We show that using multi-task learning improves recognition and more specifically, that its impact is even more significant if the amount of training data is decreasing (up to 12% relative improvement of the word error rate). All training and testing experiments are carried out on parts of the CHiME4 database.
Publication details
- DOI
- 10.1109/isspit.2017.8388673
- OpenAlex
- W2810852753
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.