Pushpak Bhattacharyya
16 papers in the PaperMetrix corpus
Papers by this author
-
Data representation methods and use of mined corpora for Indian language transliteration
2015
Our NEWS 2015 shared task submission is a PBSMT based transliteration system with the following corpus preprocessing enhancements: (i) addition of wordboundary markers, and (ii) languageindependent, overlapping character segmentation. We show that the addition of …
-
Role of Morphology Injection in SMT: A Case Study from Indian Language Perspective
2017 · arXiv (Cornell University)
Phrase-based Statistical Machine Translation (PBSMT) is commonly used for automatic translation. However, PBSMT runs into difficulty when either or both of the source and target languages are morphologically rich. Factored models are found to be …
-
Morphology Generation for Statistical Machine Translation
2017 · arXiv (Cornell University)
When translating into morphologically rich languages, Statistical MT approaches face the problem of data sparsity. The severity of the sparseness problem will be high when the corpus size of morphologically richer language is less. Even …
-
IITP at IJCNLP-2017 Task 4: Auto Analysis of Customer Feedback using CNN and GRU Network
2017 · International Joint Conference on Natural Language Processing
Analyzing customer feedback is the best way to channelize the data into new marketing strategies that benefit entrepreneurs as well as customers. Therefore an automated system which can analyze the customer behavior is in great …
-
Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER
2018
Multilingual learning for Neural Named Entity Recognition (NNER) involves jointly training a neural network for multiple languages. Typically, the goal is improving the NER performance of one of the languages (the primary language) using the …
-
Temporality as Seen through Translation: A Case Study on Hindi Texts
2017 · Edinburgh Research Explorer (University of Edinburgh)
Temporality hassignificantly contributed to various aspects of Natural Language Processing applications. In this paper, we determine the extent to which temporal orientation is preserved when a sentence is translated manually and automatically from the Hindi …
-
Extraction of Message Sequence Charts from Narrative History Text
2019
Girish Palshikar, Sachin Pawar, Sangameshwar Patil, Swapnil Hingmire, Nitin Ramrakhiyani, Harsimran Bedi, Pushpak Bhattacharyya, Vasudeva Varma. Proceedings of the First Workshop on Narrative Understanding. 2019.
-
Indic language computing
2019 · Communications of the ACM
research-article Share on Indic language computing Authors: Pushpak Bhattacharyya IIT Bombay and IIT Patna IIT Bombay and IIT PatnaView Profile , Hema Murthy IIT Madras IIT MadrasView Profile , Surangika Ranathunga University of Moratuwa University …
-
A Study of Efficacy of Cross-lingual Word Embeddings for Indian Languages
2020
Cross-lingual word embeddings have become ubiquitous for various NLP tasks. Existing literature primarily evaluate the quality of cross-lingual word embeddings on the task of Bilingual Lexicon Induction. They report very high accuracies for European languages. …
-
EmoInHindi: A Multi-label Emotion and Intensity Annotated Dataset in Hindi for Emotion Recognition in Dialogues
2022
The long-standing goal of Artificial Intelligence (AI) has been to create human-like conversational systems.Such systems should have the ability to develop an emotional connection with the users, hence emotion recognition in dialogues is an important …
-
Can Taxonomy Help? Improving Semantic Question Matching using Question\n Taxonomy
2021 · arXiv (Cornell University)
In this paper, we propose a hybrid technique for semantic question matching.\nIt uses our proposed two-layered taxonomy for English questions by augmenting\nstate-of-the-art deep learning models with question classes obtained from a\ndeep learning based question classifier. …
-
Improving Machine Translation with Phrase Pair Injection and Corpus Filtering
2023 · arXiv (Cornell University)
In this paper, we show that the combination of Phrase Pair Injection and Corpus Filtering boosts the performance of Neural Machine Translation (NMT) systems. We extract parallel phrases and sentences from the pseudo-parallel corpus and …
-
Striking a Balance between Classical and Deep Learning Approaches in Natural Language Processing Pedagogy
2024 · arXiv (Cornell University)
While deep learning approaches represent the state-of-the-art of natural language processing (NLP) today, classical algorithms and approaches still find a place in NLP textbooks and courses of recent years. This paper discusses the perspectives of …
-
One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation
2024
Tejpalsingh Siledar, Swaroop Nath, Sankara Muddu, Rupasai Rangaraju, Swaprava Nath, Pushpak Bhattacharyya, Suman Banerjee, Amey Patil, Sudhanshu Singh, Muthusamy Chelliah, Nikesh Garera. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume …
-
Brahmi-Net: A transliteration and script conversion system for languages of the Indian subcontinent
2015
We present Brahmi-Net- an online system for transliteration and script conversion for all ma-jor Indian language pairs (306 pairs). The sys-tem covers 13 Indo-Aryan languages, 4 Dra-vidian languages and English. For training the transliteration systems, …
-
Are Word Embedding-based Features Useful for Sarcasm Detection?
2016
This paper makes a simple increment to state-of-the-art in sarcasm detection research. Existing approaches are unable to capture subtle forms of context incongruity which lies at the heart of sarcasm. We explore if prior work …