preprint Open access

Wav2Letter: an End-to-End ConvNet-based Speech Recognition System

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
248
References
24
Comments
0
Paper overview

Öz

This paper presents a simple end-to-end model for speech recognition, combining a convolutional network based acoustic model and a graph decoding. It is trained to output letters, with transcribed speech, without the need for force alignment of phonemes. We introduce an automatic segmentation criterion for training from sequence annotation without alignment that is on par with CTC while being simpler. We show competitive results in word error rate on the Librispeech corpus with MFCC features, and promising results from raw waveform.

Record transparency

Publication details

DOI
10.48550/arxiv.1609.03193
OpenAlex
W2520160253
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.