Jiatao Gu
11 papers in the PaperMetrix corpus
Papers by this author
-
Neural Machine Translation with Gumbel-Greedy Decoding
2017 · arXiv (Cornell University)
Previous neural machine translation models used some heuristic search algorithms (e.g., beam search) in order to avoid solving the maximum a posteriori problem over translation sentences at test time. In this paper, we propose the …
-
Efficient Learning for Undirected Topic Models
2015
Jiatao Gu, Victor O.K. Li. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 2015.
-
Incorporating Copying Mechanism in Sequence-to-Sequence Learning
2016 · arXiv (Cornell University)
We address an important problem in sequence-to-sequence (Seq2Seq) learning referred to as copying, in which certain segments in the input sequence are selectively replicated in the output sequence. A similar phenomenon is observable in human …
-
Non-Autoregressive Neural Machine Translation
2017 · arXiv (Cornell University)
Existing approaches to neural machine translation condition each output word on previously generated outputs. We introduce a model that avoids this autoregressive property and produces its outputs in parallel, allowing an order of magnitude lower …
-
Universal Neural Machine Translation for Extremely Low Resource Languages
2018 · arXiv (Cornell University)
In this paper, we propose a new universal machine translation approach focusing on languages with a limited amount of parallel data. Our proposed approach utilizes a transfer-learning approach to share lexical and sentence level representations …
-
Search Engine Guided Neural Machine Translation
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
In this paper, we extend an attention-based neural machine translation (NMT) model by allowing it to access an entire training set of parallel sentence pairs even after training. The proposed approach consists of two stages. …
-
Meta-Learning for Low-Resource Neural Machine Translation
2018
In this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm (MAML, Finn et al., 2017) for lowresource neural machine translation (NMT). We frame low-resource translation as a metalearning problem, and we learn …
-
Levenshtein Transformer
2019 · Neural Information Processing Systems
Modern neural sequence generation models are built to either generate tokens step-by-step from scratch or (iteratively) modify a sequence of tokens bounded by a fixed length. In this work, we develop Levenshtein Transformer, a new …
-
Insertion-based Decoding with Automatically Inferred Generation Order
2019 · Transactions of the Association for Computational Linguistics
Conventional neural autoregressive decoding commonly assumes a fixed left-to-right generation order, which may be sub-optimal. In this work, we propose a novel decoding algorithm— InDIGO—which supports flexible sequence generation in arbitrary orders through insertion operations. …
-
Neural Machine Translation with Byte-Level Subwords
2020 · Proceedings of the AAAI Conference on Artificial Intelligence
Almost all existing machine translation models are built on top of character-based vocabularies: characters, subwords or words. Rare characters from noisy text or character-rich languages such as Japanese and Chinese however can unnecessarily take up …
-
Multilingual Denoising Pre-training for Neural Machine Translation
2020 · Transactions of the Association for Computational Linguistics
This paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART—a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using …