conference-paper

MoNet: A Mixture of Experts Solution for Multilingual and Low-Resource ASR Challenges

Research footprint

At a glance

Citations
3
References
38
Comments
0
Paper overview

Öz

Contemporary end-to-end automatic speech recognition models have demonstrated remarkable performance in monolingual contexts with abundant corpora. However, their efficacy is considerably compromised in multilingual and resource-scarce environments due to the models’ inability to effectively learn the diverse linguistic features from unevenly distributed corpora during training. Hence, enabling these models to assimilate a broader spectrum of effective features presents a significant challenge. In this paper, we introduce a novel multilingual automatic speech recognition network, MoNet, predicated on the Mixture of Experts (MoE) approach, which is capable of automatically and efficiently processing multilingual speech data without the need for any language-specific cues. Specifically, within MoNet, the MoE layer employs a router to select a fixed number of optimal experts from an expert pool to guide the network’s learning and provide efficacious reasoning, followed by the use of joint CTC/Attention training and the implementation of a secondary scoring mechanism to further enhance the model’s performance. We conducted hybrid training and evaluation of the proposed model using a compendium of corpora from six linguistically diverse languages with considerable variance in total training duration—Turkish, Uighur, Bashkir, Uzbek, Kyrgyz, and Russian—sourced from the publicly accessible CommonVoice dataset. Our findings indicate that compared to the baseline model, MoNet achieves significant improvements, with a maximum reduction in Word Error Rate (WER) of up to 4.1% for individual languages and an average decrease of 1.78% in WER across multiple languages, while still maintaining an inference cost comparable to that of the baseline model. The experimental outcomes underscore the significant advantages of our proposed model for low-resource, multilingual speech recognition tasks.

Record transparency

Publication details

DOI
10.1109/ijcnn60899.2024.10651105
OpenAlex
W4402352982
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.