ملف الباحث

Sneha Kudugunta

ورقتان في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation

    2020 · arXiv (Cornell University)

    Over the last few years two promising research directions in low-resource neural machine translation (NMT) have emerged. The first focuses on utilizing high-resource languages to improve the quality of low-resource languages via multilingual NMT. The …

  2. Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets

    2022 · Transactions of the Association for Computational Linguistics

    Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality …