Researcher profile

Kenneth Heafield

9 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Multi-Source Syntactic Neural Machine Translation

    2018 · arXiv (Cornell University)

    We introduce a novel multi-source technique for incorporating source syntax into neural machine translation using linearized parses. This is achieved by employing separate encoders for the sequential and parsed versions of the same source sentence; …

  2. Surprise Languages: Rapid-Response Cross-Language IR

    2019 · Edinburgh Research Explorer (University of Edinburgh)

    Sixteen years ago, the first "surprise language exercise" was conducted, in Cebuano. The evaluation goal of a surprise language exercise is to learn how well systems for a new language can be quickly built. This …

  3. Iterative Translation Refinement with Large Language Models

    2023 · arXiv (Cornell University)

    We propose iteratively prompting a large language model to self-correct a translation, with inspiration from their strong language understanding and translation capability as well as a human-like translation approach. Interestingly, multi-turn querying reduces the output's …

  4. The University of Edinburgh's Neural MT Systems for WMT17

    2017

    Rico Sennrich, Alexandra Birch, Anna Currey, Ulrich Germann, Barry Haddow, Kenneth Heafield, Antonio Valerio Miceli Barone, Philip Williams. Proceedings of the Second Conference on Machine Translation. 2017.

  5. Copied Monolingual Data Improves Low-Resource Neural Machine Translation

    2017

    We train a neural machine translation (NMT) system to both translate sourcelanguage text and copy target-language text, thereby exploiting monolingual corpora in the target language. Specifically, we create a bitext from the monolingual text in …

  6. Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering

    2018

    We posed the shared task of assigning sentence-level quality scores for a very noisy corpus of sentence pairs crawled from the web, with the goal of sub-selecting 1% and 10% of high-quality data to be …

  7. Neural Grammatical Error Correction Systems with Unsupervised Pre-training on Synthetic Data

    2019

    Considerable effort has been made to address the data sparsity problem in neural grammatical error correction. In this work, we propose a simple and surprisingly effective unsupervised synthetic error generation method based on confusion sets …

  8. ParaCrawl: Web-Scale Acquisition of Parallel Corpora

    2020

    Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, …

  9. No Language Left Behind: Scaling Human-Centered Machine Translation

    2022 · arXiv (Cornell University)

    Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset …