conference-paper

Monaural Multi-Talker Speech Recognition with Attention Mechanism and Gated Convolutional Networks

Research footprint

At a glance

Citations
23
References
39
Comments
0
Paper overview

Abstract

Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring multi-talker mixed speech sequence data corresponding to a plurality of speakers, encoding the multi-speaker mixed speech sequence data into an embedded sequence data, generating speaker specific context vectors at each frame based on the embedded sequence, generating senone posteriors for each of the speaker based on the speaker specific context vectors and updating an acoustic model by performing permutation invariant training (PIT) model training based on the senone posteriors.

Record transparency

Publication details

DOI
10.21437/interspeech.2018-1547
OpenAlex
W2889144942
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.