ملف الباحث
Sneha Kudugunta
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
2020 · arXiv (Cornell University)
Over the last few years two promising research directions in low-resource neural machine translation (NMT) have emerged. The first focuses on utilizing high-resource languages to improve the quality of low-resource languages via multilingual NMT. The …
-
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
2022 · Transactions of the Association for Computational Linguistics
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality …