Julia Kreutzer
4 papers in the PaperMetrix corpus
Papers by this author
-
Correct Me If You Can: Learning from Error Corrections and Markings
2020 · arXiv (Cornell University)
Sequence-to-sequence learning involves a trade-off between signal strength and annotation cost of training data. For example, machine translation data range from costly expert-generated translations that enable supervised learning, to weak quality-judgment feedback that facilitate reinforcement …
-
A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation
2022 · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen …
-
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
2022 · Transactions of the Association for Computational Linguistics
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality …
-
MasakhaNER: Named Entity Recognition for African Languages
2021 · Transactions of the Association for Computational Linguistics
Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition …