Conformer Based End-to-End ASR System with a New Feature Fusion
At a glance
- Citations
- 1
- References
- 22
- Comments
- 0
Abstract
In recent years, the proliferation and refinement of neural networks have engendered pioneering advancements in several domains that rely on machine learning and deep learning techniques. Neural network technology has also brought about significant advancements in the field of speech recognition. The current neural networks in various fields are dominated by transformer and conformer. This study explores the current mainstream frameworks and focuses on the research of the conformer network architecture, providing a brief description of its structure. In addition to investigating the mainstream frameworks, the authors discovered that some small plug-and-play attention mechanism frameworks introduced in previous years still yield impressive results, such as ResNet, SeNet, and others. Subsequently, the authors took a closer look at these small plug-and-play frameworks and finally selected resnet as the innovation point. Based on these two, a novel feature fusion framework, the res-Conv model, is proposed. In order to verify the effectiveness of the model, the authors applied the model to the current mainstream open-source speech recognition tool wenet, using the dataset Aishell-1 and LibriSpeech. After continuous debugging, the results show that the proposed res-Conv model, compared to the baseline, achieved a reduction in Character Error Rate (CER) of 0.71-0.82% on the Aishell dataset and a reduction in Word Error Rate (WER) of 0.47-0.53% on the LibriSpeech dataset, while almost maintaining the same number of parameters.
Publication details
- DOI
- 10.1109/prai59366.2023.10331991
- OpenAlex
- W4389302233
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.