Marc’Aurelio Ranzato
8 papers in the PaperMetrix corpus
Papers by this author
-
Efficient Lifelong Learning with A-GEM
2018 · arXiv (Cornell University)
In lifelong learning, the learner is presented with a sequence of tasks, incrementally building a data-driven prior which may be leveraged to speed up learning of a new task. In this work, we investigate the …
-
Sequence Level Training with Recurrent Neural Networks
2015 · arXiv (Cornell University)
Many natural language processing applications use language models to generate text. These models are typically trained to predict the next word in a sequence, given the previous words and some context such as an image. …
-
Word Translation Without Parallel Data
2017 · arXiv (Cornell University)
State-of-the-art methods for learning cross-lingual word embeddings have relied on bilingual dictionaries or parallel corpora. Recent studies showed that the need for parallel data supervision can be alleviated with character-level information. While these methods showed …
-
Unsupervised Machine Translation Using Monolingual Corpora Only
2017 · arXiv (Cornell University)
Machine translation has recently achieved impressive performance thanks to recent advances in deep learning and the availability of large-scale parallel corpora. There have been numerous attempts to extend these successes to low-resource language pairs, yet …
-
The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English
2019
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on …
-
Phrase-Based & Neural Unsupervised Machine Translation
2018
Machine translation systems achieve near human-level performance on some languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentences, which hinders their applicability to the majority of language pairs. This …
-
Sequence Level Training with Recurrent Neural Networks
2016 · International Conference on Learning Representations
Abstract: Many natural language processing applications use language models to generate text. These models are typically trained to predict the next word in a sequence, given the previous words and some context such as an …
-
The <scp>Flores-101</scp> Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
2022 · Transactions of the Association for Computational Linguistics
Abstract One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, …