preprint Open access

Cascaded CNN-resBiLSTM-CTC: An End-to-End Acoustic Model For Speech\n Recognition

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Automatic speech recognition (ASR) tasks are resolved by end-to-end deep\nlearning models, which benefits us by less preparation of raw data, and easier\ntransformation between languages. We propose a novel end-to-end deep learning\nmodel architecture namely cascaded CNN-resBiLSTM-CTC. In the proposed model, we\nadd residual blocks in BiLSTM layers to extract sophisticated phoneme and\nsemantic information together, and apply cascaded structure to pay more\nattention mining information of hard negative samples. By applying both simple\nFast Fourier Transform (FFT) technique and n-gram language model (LM) rescoring\nmethod, we manage to achieve word error rate (WER) of 3.41% on LibriSpeech test\nclean corpora. Furthermore, we propose a new batch-varied method to speed up\nthe training process in length-varied tasks, which result in 25% less training\ntime.\n

Record transparency

Publication details

DOI
10.48550/arxiv.1810.12001
OpenAlex
W4289362596
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.