preprint Open access

Strategies for Language Identification in Code-Mixed Low Resource Languages

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
0
References
22
Comments
0
Paper overview

Öz

In recent years, substantial work has been done on language tagging of code-mixed data, but most of them use large amounts of data to build their models. In this article, we present three strategies to build a word level language tagger for code-mixed data using very low resources. Each of them secured an accuracy higher than our baseline model, and the best performing system got an accuracy around 91%. Combining all, the ensemble system achieved an accuracy of around 92.6%.

Record transparency

Publication details

DOI
10.48550/arxiv.1810.07156
OpenAlex
W2896463882
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.