conference-paper

A feature lightweight method in optimized acoustic encoder

Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

This paper is based on end-to-end speech recognition based on convolutional neural network technology. The problem that convolutional neural networks are difficult to balance accuracy and model size is analysed and studied. A new acoustic encoder is proposed to optimize the extraction of speech features. The effectiveness of the proposed method is verified by an end-to-end speech recognition model. The new acoustic encoder focuses the Global context information with the local information obtained by convolution, Global context information is added. The convolution depth is increased while the convolution kernel size is reduced. It also reduces the amount of parameter calculation. At the same time, RNN-Transducer is used as the model architecture to jointly optimize acoustic features and text features, in order to obtain more effective context information. On the LibriSpeech dataset, the improved acoustic encoder achieves 9.83% word error rate and 4.61% sentence error rate with only 7M model. Compared with the baseline model of convolutional neural network, the word error rate is relatively reduced by 0.52%, and the sentence error rate is relatively reduced by 1.22%.

Record transparency

Publication details

DOI
10.1117/12.2681644
OpenAlex
W4379013316
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.