End-To-End Lingao Dialect Speech Recognition
At a glance
- Citations
- 0
- References
- 14
- Comments
- 0
Abstract
The appearance of deep learning and large-scale language model based on Transformer brings new opportunities and challenges to speech recognition. Although Transformer-based end-to-end speech recognition technology uses a multi-attention mechanism to model the global context, which can not only improve the accuracy of speech recognition, but also greatly speed up the training speed of the model through parallel computing. However, the Transformer model lacks the ability of extracting fine features, so in order to improve the recognition effect of the model, the end-to-end speech recognition technology is adopted in the research of Lingao dialect speech recognition, conformer is used instead of Transformer at the coding end, CTC decoder and attention decoder are combined to decode. The experimental results show that the method is effective and the model character error rate is as low as 8.04% on the test set.
Publication details
- DOI
- 10.1109/acait63902.2024.11022297
- OpenAlex
- W4411173697
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.