ملف الباحث

Sunayana Sitaram

8 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. A Survey of Code-switched Speech and Language Processing

    2019 · arXiv (Cornell University)

    Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched Speech and Natural Language Processing. …

  2. Predicting the Performance of Multilingual NLP Models

    2021 · arXiv (Cornell University)

    Recent advancements in NLP have given us models like mBERT and XLMR that can serve over 100 languages. The languages that these models are evaluated on, however, are very few in number, and it is …

  3. On the Calibration of Massively Multilingual Language Models

    2022

    Massively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprising effectiveness in cross-lingual transfer. While there has been much work in evaluating these models for their performance on a variety of tasks …

  4. MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations

    2024 · arXiv (Cornell University)

    In recent years, self-supervised pre-training methods have gained significant traction in learning high-level information from raw speech. Among these methods, HuBERT has demonstrated SOTA performance in automatic speech recognition (ASR). However, HuBERT's performance lags behind …

  5. MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models

    2024

    Parameter Efficient Finetuning (PEFT) has emerged as a viable solution for improving the performance of Large Language Models (LLMs) without requiring massive resources and compute.Prior work on multilingual evaluation has shown that there is a …

  6. Multilingual CheckList: Generation and Evaluation

    2022

    Multilingual evaluation benchmarks usually contain limited high-resource languages and do not test models for specific linguistic capabilities.CheckList (Ribeiro et al., 2020) is a template-based evaluation approach that tests models for specific capabilities.The CheckList template creation …

  7. Polyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation Learning

    2016

    Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: …

  8. Language Modeling for Code-Mixing: The Role of Linguistic Theory based Synthetic Data

    2018

    Adithya Pratapa, Gayatri Bhat, Monojit Choudhury, Sunayana Sitaram, Sandipan Dandapat, Kalika Bali. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.