Self-attention Networks for Connectionist Temporal Classification in Speech Recognition
At a glance
- Citations
- 133
- References
- 66
- Comments
- 0
Abstract
The success of self-attention in NLP has led to recent applications in end-to-end encoder-decoder architectures for speech recognition. Separately, connectionist temporal classification (CTC) has matured as an alignment-free, non-autoregressive approach to sequence transduction, either by itself or in various multitask and decoding frameworks. We propose SAN-CTC, a deep, fully self-attentional network for CTC, and show it is tractable and competitive for end-to-end speech recognition. SAN-CTC trains quickly and outperforms existing CTC models and most encoder-decoder models, with character error rates (CERs) of 4.7% in 1 day on WSJ eval92 and 2.8% in 1 week on LibriSpeech test-clean, with a fixed architecture and one GPU. Similar improvements hold for WERs after LM decoding. We motivate the architecture for speech, evaluate position and down-sampling approaches, and explore how label alphabets (character, phoneme, subword) affect attention heads and performance.
Publication details
- DOI
- 10.1109/icassp.2019.8682539
- OpenAlex
- W2911291251
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.