conference-paper
Word Level Language Identification and Back-Transliteration
Research footprint
At a glance
- Citations
- 2
- References
- 18
- Comments
- 0
Paper overview
Abstract
In this paper, I describe a Rule based and List-Searching system for Word-Level Language Identification and Named Entity recognition in bilingual text. My method uses dictionary search, rules for LI and CRF++, character n-gram for NER. The model does not use any Language specific rules, therefore can easily be replicated on most languages having mixed pair with English. The model also does back-transliteration of words into native language script and recognizes named entity. The model performance is carried on the test sets provided by the shared task on language Identification for English Hindi (En-Hi) Pair, Microsoft Research India. The experimental results show a consistent performance with high precision.
Record transparency
Publication details
- DOI
- 10.1145/2824864.2824884
- OpenAlex
- W2176101793
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.