conference-paper

Word Level Language Identification and Back-Transliteration

Research footprint

At a glance

Citations
2
References
18
Comments
0
Paper overview

Abstract

In this paper, I describe a Rule based and List-Searching system for Word-Level Language Identification and Named Entity recognition in bilingual text. My method uses dictionary search, rules for LI and CRF++, character n-gram for NER. The model does not use any Language specific rules, therefore can easily be replicated on most languages having mixed pair with English. The model also does back-transliteration of words into native language script and recognizes named entity. The model performance is carried on the test sets provided by the shared task on language Identification for English Hindi (En-Hi) Pair, Microsoft Research India. The experimental results show a consistent performance with high precision.

Record transparency

Publication details

DOI
10.1145/2824864.2824884
OpenAlex
W2176101793
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.