Researcher profile

Jiatao Gu

11 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Neural Machine Translation with Gumbel-Greedy Decoding

    2017 · arXiv (Cornell University)

    Previous neural machine translation models used some heuristic search algorithms (e.g., beam search) in order to avoid solving the maximum a posteriori problem over translation sentences at test time. In this paper, we propose the …

  2. Efficient Learning for Undirected Topic Models

    2015

    Jiatao Gu, Victor O.K. Li. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 2015.

  3. Incorporating Copying Mechanism in Sequence-to-Sequence Learning

    2016 · arXiv (Cornell University)

    We address an important problem in sequence-to-sequence (Seq2Seq) learning referred to as copying, in which certain segments in the input sequence are selectively replicated in the output sequence. A similar phenomenon is observable in human …

  4. Non-Autoregressive Neural Machine Translation

    2017 · arXiv (Cornell University)

    Existing approaches to neural machine translation condition each output word on previously generated outputs. We introduce a model that avoids this autoregressive property and produces its outputs in parallel, allowing an order of magnitude lower …

  5. Universal Neural Machine Translation for Extremely Low Resource Languages

    2018 · arXiv (Cornell University)

    In this paper, we propose a new universal machine translation approach focusing on languages with a limited amount of parallel data. Our proposed approach utilizes a transfer-learning approach to share lexical and sentence level representations …

  6. Search Engine Guided Neural Machine Translation

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    In this paper, we extend an attention-based neural machine translation (NMT) model by allowing it to access an entire training set of parallel sentence pairs even after training. The proposed approach consists of two stages. …

  7. Meta-Learning for Low-Resource Neural Machine Translation

    2018

    In this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm (MAML, Finn et al., 2017) for lowresource neural machine translation (NMT). We frame low-resource translation as a metalearning problem, and we learn …

  8. Levenshtein Transformer

    2019 · Neural Information Processing Systems

    Modern neural sequence generation models are built to either generate tokens step-by-step from scratch or (iteratively) modify a sequence of tokens bounded by a fixed length. In this work, we develop Levenshtein Transformer, a new …

  9. Insertion-based Decoding with Automatically Inferred Generation Order

    2019 · Transactions of the Association for Computational Linguistics

    Conventional neural autoregressive decoding commonly assumes a fixed left-to-right generation order, which may be sub-optimal. In this work, we propose a novel decoding algorithm— InDIGO—which supports flexible sequence generation in arbitrary orders through insertion operations. …

  10. Neural Machine Translation with Byte-Level Subwords

    2020 · Proceedings of the AAAI Conference on Artificial Intelligence

    Almost all existing machine translation models are built on top of character-based vocabularies: characters, subwords or words. Rare characters from noisy text or character-rich languages such as Japanese and Chinese however can unnecessarily take up …

  11. Multilingual Denoising Pre-training for Neural Machine Translation

    2020 · Transactions of the Association for Computational Linguistics

    This paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART—a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using …