Kenton Murray
4 papers in the PaperMetrix corpus
Papers by this author
-
BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation
2021 · arXiv (Cornell University)
The success of bidirectional encoders using masked language models, such as BERT, on numerous natural language processing tasks has prompted researchers to attempt to incorporate these pre-trained models into neural machine translation (NMT) systems. However, …
-
NL-Augmenter 🦎 → 🐍 A Framework for Task-Sensitive Natural Language Augmentation
2023 · Northern European Journal of Language Technology
Data augmentation is an important method for evaluating the robustness of and enhancing the diversity of training data for natural language processing (NLP) models. In this paper, we present NL-Augmenter, a new participatory Python-based natural …
-
MegaWika: Millions of reports and their sources across 50 diverse languages
2023 · arXiv (Cornell University)
To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced source materials. We process …
-
Evaluating Large Language Models along Dimensions of Language Variation: A Systematik Invesdigatiom uv Cross-lingual Generalization
2024 · arXiv (Cornell University)
While large language models exhibit certain cross-lingual generalization capabilities, they suffer from performance degradation (PD) on unseen closely-related languages (CRLs) and dialects relative to their high-resource language neighbour (HRLN). However, we currently lack a fundamental …