ملف الباحث

Anoop Kunchukuttan

9 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Data representation methods and use of mined corpora for Indian language transliteration

    2015

    Our NEWS 2015 shared task submission is a PBSMT based transliteration system with the following corpus preprocessing enhancements: (i) addition of wordboundary markers, and (ii) languageindependent, overlapping character segmentation. We show that the addition of …

  2. Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER

    2018

    Multilingual learning for Neural Named Entity Recognition (NNER) involves jointly training a neural network for multiple languages. Typically, the goal is improving the NER performance of one of the languages (the primary language) using the …

  3. A Brief Survey of Multilingual Neural Machine Translation

    2019 · arXiv (Cornell University)

    We present a survey on multilingual neural machine translation (MNMT), which has gained a lot of traction in the recent years. MNMT has been useful in improving translation quality as a result of knowledge transfer. …

  4. Evaluating Inter-Bilingual Semantic Parsing for Indian Languages

    2023 · arXiv (Cornell University)

    Despite significant progress in Natural Language Generation for Indian languages (IndicNLP), there is a lack of datasets around complex structured tasks such as semantic parsing. One reason for this imminent gap is the complexity of …

  5. Bhasha-Abhijnaanam: Native-script and romanized Language Identification for 22 Indic languages

    2023 · arXiv (Cornell University)

    We create publicly available language identification (LID) datasets and models in all 22 Indian languages listed in the Indian constitution in both native-script and romanized text. First, we create Bhasha-Abhijnaanam, a language identification test set …

  6. Brahmi-Net: A transliteration and script conversion system for languages of the Indian subcontinent

    2015

    We present Brahmi-Net- an online system for transliteration and script conversion for all ma-jor Indian language pairs (306 pairs). The sys-tem covers 13 Indo-Aryan languages, 4 Dra-vidian languages and English. For training the transliteration systems, …

  7. Overview of the 8th Workshop on Asian Translation

    2021

    Toshiaki Nakazawa, Hideki Nakayama, Chenchen Ding, Raj Dabre, Shohei Higashiyama, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Shantipriya Parida, Ondřej Bojar, Chenhui Chu, Akiko Eriguchi, Kaori Abe, Yusuke Oda, Sadao Kurohashi. Proceedings of …

  8. Overview of the 7th Workshop on Asian Translation

    2020

    Toshiaki Nakazawa, Hideki Nakayama, Chenchen Ding, Raj Dabre, Shohei Higashiyama, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Shantipriya Parida, Ondřej Bojar, Sadao Kurohashi. Proceedings of the 7th Workshop on Asian Translation. 2020.

  9. IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages

    2020

    In this paper, we introduce NLP resources for 11 major Indian languages from two major language families. These resources include: (a) large-scale sentence-level monolingual corpora, (b) pre-trained word embeddings, (c) pre-trained language models, and (d) …