Anoop Kunchukuttan
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Data representation methods and use of mined corpora for Indian language transliteration
2015
Our NEWS 2015 shared task submission is a PBSMT based transliteration system with the following corpus preprocessing enhancements: (i) addition of wordboundary markers, and (ii) languageindependent, overlapping character segmentation. We show that the addition of …
-
Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER
2018
Multilingual learning for Neural Named Entity Recognition (NNER) involves jointly training a neural network for multiple languages. Typically, the goal is improving the NER performance of one of the languages (the primary language) using the …
-
A Brief Survey of Multilingual Neural Machine Translation
2019 · arXiv (Cornell University)
We present a survey on multilingual neural machine translation (MNMT), which has gained a lot of traction in the recent years. MNMT has been useful in improving translation quality as a result of knowledge transfer. …
-
Evaluating Inter-Bilingual Semantic Parsing for Indian Languages
2023 · arXiv (Cornell University)
Despite significant progress in Natural Language Generation for Indian languages (IndicNLP), there is a lack of datasets around complex structured tasks such as semantic parsing. One reason for this imminent gap is the complexity of …
-
Bhasha-Abhijnaanam: Native-script and romanized Language Identification for 22 Indic languages
2023 · arXiv (Cornell University)
We create publicly available language identification (LID) datasets and models in all 22 Indian languages listed in the Indian constitution in both native-script and romanized text. First, we create Bhasha-Abhijnaanam, a language identification test set …
-
Brahmi-Net: A transliteration and script conversion system for languages of the Indian subcontinent
2015
We present Brahmi-Net- an online system for transliteration and script conversion for all ma-jor Indian language pairs (306 pairs). The sys-tem covers 13 Indo-Aryan languages, 4 Dra-vidian languages and English. For training the transliteration systems, …
-
Overview of the 8th Workshop on Asian Translation
2021
Toshiaki Nakazawa, Hideki Nakayama, Chenchen Ding, Raj Dabre, Shohei Higashiyama, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Shantipriya Parida, Ondřej Bojar, Chenhui Chu, Akiko Eriguchi, Kaori Abe, Yusuke Oda, Sadao Kurohashi. Proceedings of …
-
Overview of the 7th Workshop on Asian Translation
2020
Toshiaki Nakazawa, Hideki Nakayama, Chenchen Ding, Raj Dabre, Shohei Higashiyama, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Shantipriya Parida, Ondřej Bojar, Sadao Kurohashi. Proceedings of the 7th Workshop on Asian Translation. 2020.
-
IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages
2020
In this paper, we introduce NLP resources for 11 major Indian languages from two major language families. These resources include: (a) large-scale sentence-level monolingual corpora, (b) pre-trained word embeddings, (c) pre-trained language models, and (d) …