Researcher profile

Graham Neubig

45 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Pseudogen: A Tool to Automatically Generate Pseudo-Code from Source Code

    2015

    Understanding the behavior of source code written in an unfamiliar programming language is difficult. One way to aid understanding of difficult code is to add corresponding pseudo-code, which describes in detail the workings of the …

  2. Generalizing and Hybridizing Count-based and Neural Language Models

    2016

    Language models (LMs) are statistical models that calculate probabilities over sequences of words or other discrete symbols. Currently two major paradigms for language modeling exist: count-based n-gram models, which have advantages of scalability and test-time …

  3. A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    Beam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do …

  4. Neural Machine Translation Models using Binarized Prediction and Error Correction

    2018 · Journal of Natural Language Processing

    本論文では,ニューラル翻訳モデルで問題となる出力層の時間・空間計算量を,二値符号を用いた予測法により大幅に削減する手法を提案する.提案手法では従来のソフトマックスのように各単語のスコアを直接求めるのではなく,各単語に対応付けられたビット列を予測することにより,間接的に出力単語の確率を求める.これにより,最も効率的な場合で従来法の対数程度まで出力層の計算量を削減可能である.このようなモデルはソフトマックスよりも推定が難しく,単体で適用した場合には翻訳精度の低下を招く.このため,本研究では提案手法の性能を補償するために,従来法との混合モデル,および二値符号に対する誤り訂正手法の適用という 2 点の改良も提案する.日英・英日翻訳タスクを用いた評価実験により,提案法が従来法と比較して同等程度の BLEU を達成可能であるとともに,出力層に要するメモリを数十分の 1 に削減し,CPU での実行速度を 5 倍から 10 倍程度に向上可能であることを示す.

  5. Learning to Describe Phrases with Local and Global Contexts

    2018 · arXiv (Cornell University)

    When reading a text, it is common to become stuck on unfamiliar words and phrases, such as polysemous words with novel senses, rarely used idioms, internet slang, or emerging entities. If we humans cannot figure …

  6. Multilingual Neural Machine Translation With Soft Decoupled Encoding

    2019 · arXiv (Cornell University)

    Multilingual training of neural machine translation (NMT) systems has led to impressive accuracy improvements on low-resource languages. However, there are still significant challenges in efficiently learning word representations in the face of paucity of data. …

  7. An Adversarial Approach to High-Quality, Sentiment-Controlled Neural Dialogue Generation

    2019 · arXiv (Cornell University)

    In this work, we propose a method for neural dialogue response generation that allows not only generating semantically reasonable responses according to the dialogue history, but also explicitly controlling the sentiment of the response via …

  8. Improving Robustness of Machine Translation with Synthetic Noise

    2019

    Vaibhav Vaibhav, Sumeet Singh, Craig Stewart, Graham Neubig. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.

  9. Syntactic Matching Methods in Pivot Translation

    2018 · Journal of Natural Language Processing

    The pivot translation is useful method for translating between languages that contain little or no parallel data by utilizing equivalents in an intermediate language such as English. Commonly, phrase-based or tree-based pivot translation methods merge …

  10. compare-mt: A Tool for Holistic Comparison of Language Generation Systems

    2019

    Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, Xinyi Wang. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). 2019.

  11. Choosing Transfer Languages for Cross-Lingual Learning

    2019 · arXiv (Cornell University)

    Cross-lingual transfer, where a high-resource transfer language is used to improve the accuracy of a low-resource task language, is now an invaluable tool for improving performance of natural language processing (NLP) on low-resource languages. However, …

  12. Improving Neural Machine Translation through Phrase-based Forced Decoding

    2017 · International Joint Conference on Natural Language Processing

    Compared to traditional statistical machine translation (SMT), neural machine translation (NMT) often sacrifices adequacy for the sake of fluency. We propose a method to combine the advantages of traditional SMT and NMT by exploiting an …

  13. TICO-19: the Translation Initiative for COvid-19

    2020

    Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, Sylwia …

  14. OCR Post Correction for Endangered Language Texts

    2020

    There is little to no data available to build natural language processing models for most endangered languages. However, textual data in these languages often exists in formats that are not machine-readable, such as paper books …

  15. Automatic Interlinear Glossing for Under-Resourced Languages Leveraging Translations

    2020

    Interlinear Glossed Text (IGT) is a widely used format for encoding linguistic information in language documentation projects and scholarly papers. Manual production of IGT takes time and requires linguistic expertise. We attempt to address this …

  16. ExplainaBoard: An Explainable Leaderboard for NLP

    2021 · arXiv (Cornell University)

    With the rapid development of NLP research, leaderboards have emerged as one tool to track the performance of various systems on various NLP tasks. They are effective in this goal to some extent, but generally …

  17. How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering

    2020 · arXiv (Cornell University)

    Recent works have shown that language models (LM) capture different types of knowledge regarding facts or common sense. However, because no model is perfect, they still fail to provide appropriate answers in many cases. In …

  18. When is Wall a Pared and when a Muro? -- Extracting Rules Governing Lexical Selection

    2021 · arXiv (Cornell University)

    Learning fine-grained distinctions between vocabulary items is a key challenge in learning a new language. For example, the noun "wall" has different lexical manifestations in Spanish -- "pared" refers to an indoor wall while "muro" …

  19. Distributionally Robust Models with Parametric Likelihood Ratios

    2022 · arXiv (Cornell University)

    As machine learning models are deployed ever more broadly, it becomes increasingly important that they are not only able to perform well on their training distribution, but also yield accurate predictions when confronted with distribution …

  20. AUTOLEX: An Automatic Framework for Linguistic Exploration

    2022 · arXiv (Cornell University)

    Each language has its own complex systems of word, phrase, and sentence construction, the guiding principles of which are often summarized in grammar descriptions for the consumption of linguists or language learners. However, manual creation …

  21. Systematic Inequalities in Language Technology Performance across the World's Languages

    2021 · arXiv (Cornell University)

    Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development. While the performance of NLP methods has grown enormously over the …

  22. NusaCrowd: Open Source Initiative for Indonesian NLP Resources

    2022 · arXiv (Cornell University)

    We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data …

  23. Learning to Generate Pseudo-Code from Source Code Using Statistical Machine Translation

    2015

    Pseudo-code written in natural language can aid the comprehension of source code in unfamiliar programming languages. However, the great majority of source code has no corresponding pseudo-code, because pseudo-code is redundant and laborious to create. …

  24. Selecting Syntactic, Non-redundant Segments in Active Learning for Machine Translation

    2016

    Akiva Miura, Graham Neubig, Michael Paul, Satoshi Nakamura. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.

  25. What Do Recurrent Neural Network Grammars Learn About Syntax?

    2017

    Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.

  26. A Syntactic Neural Model for General-Purpose Code Generation

    2017 · arXiv (Cornell University)

    We consider the problem of parsing natural language descriptions into source code written in a general-purpose programming language like Python. Existing data-driven methods treat this problem as a language generation task without considering the underlying …

  27. Cross-Lingual Word Embeddings for Low-Resource Language Modeling

    2017

    Oliver Adams, Adam Makarucha, Graham Neubig, Steven Bird, Trevor Cohn. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.

  28. XNMT: The eXtensible Neural Machine Translation Toolkit

    2018 · arXiv (Cornell University)

    This paper describes XNMT, the eXtensible Neural Machine Translation toolkit. XNMT distin- guishes itself from other open-source NMT toolkits by its focus on modular code design, with the purpose of enabling fast iteration in research …

  29. When and Why are Pre-trained Word Embeddings Useful for Neural Machine Translation?

    2018 · arXiv (Cornell University)

    The performance of Neural Machine Translation (NMT) systems often suffers in low-resource scenarios where sufficiently large-scale parallel corpora cannot be obtained. Pre-trained word embeddings have proven to be invaluable for improving performance in natural language …

  30. Rapid Adaptation of Neural Machine Translation to New Languages

    2018

    This paper examines the problem of adapting neural machine translation systems to new, low-resourced languages (LRLs) as effectively and rapidly as possible. We propose methods based on starting with massively multilingual "seed models", which can …

  31. SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation

    2018

    In this work, we examine methods for data augmentation for text-based tasks such as neural machine translation (NMT). We formulate the design of a data augmentation policy with desirable properties as an optimization problem, and …

  32. TRANX: A Transition-based Neural Abstract Syntax Parser for Semantic Parsing and Code Generation

    2018

    We present TRANX, a transition-based neural semantic parser that maps natural language (NL) utterances into formal meaning representations (MRs). TRANX uses a transition system based on the abstract syntax description language for the target MR, …

  33. On Evaluation of Adversarial Perturbations for Sequence-to-Sequence Models

    2019

    Paul Michel, Xian Li, Graham Neubig, Juan Pino. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.

  34. Competence-based Curriculum Learning for Neural Machine Translation

    2019

    Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabas Poczos, Tom Mitchell. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short …

  35. Incorporating Discrete Translation Lexicons into Neural Machine Translation

    2016

    Neural machine translation (NMT) often makes mistakes in translating low-frequency content words that are essential to understanding the meaning of the sentence. We propose a method to alleviate this problem by augmenting NMT systems with …

  36. Controlling Output Length in Neural Encoder-Decoders

    2016

    Neural encoder-decoder models have shown great success in many sequence generation tasks. However, previous work has not investigated situations in which we would like to control the length of encoder-decoder outputs. This capability is crucial …

  37. Guiding Neural Machine Translation with Retrieved Translation Pieces

    2018

    Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, Satoshi Nakamura. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.

  38. A Probabilistic Formulation of Unsupervised Text Style Transfer

    2020 · arXiv (Cornell University)

    We present a deep generative model for unsupervised text style transfer that unifies previously proposed non-generative techniques. Our probabilistic approach models non-parallel data from two domains as a partially observed parallel corpus. By hypothesizing a …

  39. XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

    2020 · arXiv (Cornell University)

    Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, …

  40. Interpretable Multi-dataset Evaluation for Named Entity Recognition

    2020

    With the proliferation of models for natural language processing tasks, it is even harder to understand the differences between models and their relative merits. Simply looking at differences between holistic metrics such as accuracy, BLEU, …

  41. BARTScore: Evaluating Generated Text as Text Generation

    2021 · arXiv (Cornell University)

    A wide variety of NLP applications, such as machine translation, summarization, and dialog, involve text generation. One major challenge for these applications is how to evaluate whether such generated texts are actually fluent, accurate, or …

  42. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

    2022 · ACM Computing Surveys

    This article surveys and organizes research works in a new paradigm in natural language processing, which we dub “prompt-based learning.” Unlike traditional supervised learning, which trains a model to take in an input x and …

  43. MasakhaNER: Named Entity Recognition for African Languages

    2021 · Transactions of the Association for Computational Linguistics

    Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition …

  44. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

    2021 · arXiv (Cornell University)

    This paper surveys and organizes research works in a new paradigm in natural language processing, which we dub "prompt-based learning". Unlike traditional supervised learning, which trains a model to take in an input x and …

  45. Active Retrieval Augmented Generation

    2023

    Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.