preprint Open access

End-to-End Code Switching Language Models for Automatic Speech Recognition

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
0
References
18
Comments
0
Paper overview

Abstract

In this paper, we particularly work on the code-switched text, one of the most common occurrences in the bilingual communities across the world. Due to the discrepancies in the extraction of code-switched text from an Automated Speech Recognition(ASR) module, and thereby extracting the monolingual text from the code-switched text, we propose an approach for extracting monolingual text using Deep Bi-directional Language Models(LM) such as BERT and other Machine Translation models, and also explore different ways of extracting code-switched text from the ASR model. We also explain the robustness of the model by comparing the results of Perplexity and other different metrics like WER, to the standard bi-lingual text output without any external information.

Record transparency

Publication details

DOI
10.48550/arxiv.2006.08870
OpenAlex
W3035687429
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.