Multi-task learning deep neural networks for speech feature denoising
At a glance
- Citations
- 8
- References
- 11
- Comments
- 0
Öz
Traditional automatic speech recognition (ASR) systems usually\nget a sharp performance drop when noise presents in\nspeech. To make a robust ASR, we introduce a new model using\nthe multi-task learning deep neural networks (MTL-DNN)\nto solve the speech denoising task in feature level. In this model,\nthe networks are initialized by pre-training restricted Boltzmann\nmachines (RBM) and fine-tuned by jointly learning multiple\ninteractive tasks using a shared representation. In multi-task\nlearning, we choose a noisy-clean speech pair fitting task as the\nprimary task and separately explore two constraints as the secondary\ntasks: phone label and phone cluster. In experiments,\nthe denoised speech is reconstructed by the MTL-DNN using\nthe noisy speech as input and it is respectively evaluated by the\nDNN-hidden Markov model (HMM) based and the Gaussian\nMixture Model (GMM)-HMM based ASR systems. Results\nshow that, using the denoised speech, the word error rate (WER)\nis respectively reduced by 53.14% and 34.84% compared with\nbaselines. The MTL-DNN model also outperforms the general\nsingle-task learning deep neural networks (STL-DNN) model\nwith a performance improvement of 4.93% and 3.88% respectively.
Publication details
- DOI
- 10.21437/interspeech.2015-532
- OpenAlex
- W2396990910
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.