conference-paper

Conformer Based End-to-End ASR System with a New Feature Fusion

Research footprint

At a glance

Citations
1
References
22
Comments
0
Paper overview

Abstract

In recent years, the proliferation and refinement of neural networks have engendered pioneering advancements in several domains that rely on machine learning and deep learning techniques. Neural network technology has also brought about significant advancements in the field of speech recognition. The current neural networks in various fields are dominated by transformer and conformer. This study explores the current mainstream frameworks and focuses on the research of the conformer network architecture, providing a brief description of its structure. In addition to investigating the mainstream frameworks, the authors discovered that some small plug-and-play attention mechanism frameworks introduced in previous years still yield impressive results, such as ResNet, SeNet, and others. Subsequently, the authors took a closer look at these small plug-and-play frameworks and finally selected resnet as the innovation point. Based on these two, a novel feature fusion framework, the res-Conv model, is proposed. In order to verify the effectiveness of the model, the authors applied the model to the current mainstream open-source speech recognition tool wenet, using the dataset Aishell-1 and LibriSpeech. After continuous debugging, the results show that the proposed res-Conv model, compared to the baseline, achieved a reduction in Character Error Rate (CER) of 0.71-0.82% on the Aishell dataset and a reduction in Word Error Rate (WER) of 0.47-0.53% on the LibriSpeech dataset, while almost maintaining the same number of parameters.

Record transparency

Publication details

DOI
10.1109/prai59366.2023.10331991
OpenAlex
W4389302233
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.