conference-paper
Monaural Multi-Talker Speech Recognition with Attention Mechanism and Gated Convolutional Networks
Research footprint
At a glance
- Citations
- 23
- References
- 39
- Comments
- 0
Paper overview
Abstract
Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring multi-talker mixed speech sequence data corresponding to a plurality of speakers, encoding the multi-speaker mixed speech sequence data into an embedded sequence data, generating speaker specific context vectors at each frame based on the embedded sequence, generating senone posteriors for each of the speaker based on the speaker specific context vectors and updating an acoustic model by performing permutation invariant training (PIT) model training based on the senone posteriors.
Record transparency
Publication details
- DOI
- 10.21437/interspeech.2018-1547
- OpenAlex
- W2889144942
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.