Researcher profile

Philipp Koehn

19 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. The Operation Sequence Model—Combining N-Gram-Based and Phrase-Based Statistical Machine Translation

    2015 · Computational Linguistics

    In this article, we present a novel machine translation model, the Operation Sequence Model (OSM), which combines the benefits of phrase-based and N-gram-based statistical machine translation (SMT) and remedies their drawbacks. The model represents the …

  2. Knowledge Tracing in Sequential Learning of Inflected Vocabulary

    2017

    We present a feature-rich knowledge tracing method that captures a student's acquisition and retention of knowledge during a foreign language phrase learning task. We model the student's behavior as making predictions under a log-linear model, …

  3. Findings of the 2018 Conference on Machine Translation (WMT18)

    2018

    Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Christof Monz. Proceedings of the Third Conference on Machine Translation: Shared Task Papers. 2018.

  4. A Massive Collection of Cross-Lingual Web-Document Pairs.

    2019 · arXiv (Cornell University)

    Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. Small-scale efforts have been made to collect aligned document level data on …

  5. TICO-19: the Translation Initiative for COvid-19

    2020

    Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, Sylwia …

  6. Evaluating Saliency Methods for Neural Language Models

    2021

    Saliency methods are widely used to interpret neural network predictions, but different variants of saliency methods often disagree even on the interpretations of the same prediction made by the same model. In these cases, how …

  7. Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation

    2024

    Longyue Wang, Siyou Liu, Chenyang Lyu, Wenxiang Jiao, Xing Wang, Jiahao Xu, Zhaopeng Tu, Yan Gu, Weiyu Chen, Minghao Wu, Liting Zhou, Philipp Koehn, Andy Way, Yulin Yuan. Proceedings of the Ninth Conference on Machine …

  8. Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents

    2025 · arXiv (Cornell University)

    Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To train agents for this dynamic process, particularly in multi-modal contexts, we introduce …

  9. Findings of the 2015 Workshop on Statistical Machine Translation

    2015

    Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, Marco Turchi. Proceedings of the Tenth Workshop on Statistical …

  10. Findings of the 2016 Conference on Machine Translation

    2016

    Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, …

  11. Predicting Target Language CCG Supertags Improves Neural Machine Translation

    2017

    Neural machine translation (NMT) models are able to partially learn syntactic information from sequential lexical information. Still, some complex syntactic phenomena such as prepositional phrase attachment are poorly modeled. This work aims to answer two …

  12. Findings of the 2017 Conference on Machine Translation (WMT17)

    2017

    Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, Marco Turchi. Proceedings of the …

  13. Iterative Back-Translation for Neural Machine Translation

    2018

    We present iterative back-translation, a method for generating increasingly better synthetic parallel data from monolingual data to train neural machine translation systems. Our proposed method is very simple yet effective and highly applicable in practice. …

  14. Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering

    2018

    We posed the shared task of assigning sentence-level quality scores for a very noisy corpus of sentence pairs crawled from the web, with the goal of sub-selecting 1% and 10% of high-quality data to be …

  15. The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English

    2019

    Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on …

  16. Findings of the 2019 Conference on Machine Translation (WMT19)

    2019

    Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, Marcos Zampieri. Proceedings of the Fourth …

  17. Saliency-driven Word Alignment Interpretation for Neural Machine Translation

    2019

    Despite their original goal to jointly learn to align and translate, Neural Machine Translation (NMT) models, especially Transformer, are often perceived as not learning interpretable word alignments. In this paper, we show that NMT models …

  18. ParaCrawl: Web-Scale Acquisition of Parallel Corpora

    2020

    Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, …

  19. No Language Left Behind: Scaling Human-Centered Machine Translation

    2022 · arXiv (Cornell University)

    Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset …