Researcher profile

Vishrav Chaudhary

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. A Massive Collection of Cross-Lingual Web-Document Pairs.

    2019 · arXiv (Cornell University)

    Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. Small-scale efforts have been made to collect aligned document level data on …

  2. The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English

    2019

    Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on …

  3. WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia

    2021

    Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, Francisco Guzmán. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.

  4. Unsupervised Cross-lingual Representation Learning at Scale

    2020

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.

  5. Beyond English-Centric Multilingual Machine Translation

    2020 · arXiv (Cornell University)

    Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only …

  6. The <scp>Flores-101</scp> Evaluation Benchmark for Low-Resource and Multilingual Machine Translation

    2022 · Transactions of the Association for Computational Linguistics

    Abstract One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, …