Minh-Thang Luong
11 papers in the PaperMetrix corpus
Papers by this author
-
Multi-task Sequence to Sequence Learning
2015 · arXiv (Cornell University)
Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. …
-
Achieving Open Vocabulary Neural Machine Translation with Hybrid Word-Character Models
2016
Nearly all previous work on neural machine translation (NMT) has used quite restricted vocabularies, perhaps with a subsequent method to patch in unknown words. This paper presents a novel wordcharacter solution to achieving open vocabulary …
-
Massive Exploration of Neural Machine Translation Architectures
2017 · arXiv (Cornell University)
Neural Machine Translation (NMT) has shown remarkable progress over the past few years, with production systems now being deployed to end-users. As the field is moving rapidly, it has become unclear which elements of NMT …
-
Online and Linear-Time Attention by Enforcing Monotonic Alignments
2017 · arXiv (Cornell University)
Recurrent neural network models with an attention mechanism have proven to be extremely effective on a wide variety of sequence-to-sequence problems. However, the fact that soft attention mechanisms perform a pass over the entire input …
-
QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
2018 · arXiv (Cornell University)
Current end-to-end machine reading and question answering (Q\&A) models are primarily based on recurrent neural networks (RNNs) with attention. Despite their success, these models are often slow for both training and inference due to the …
-
Semi-Supervised Sequence Modeling with Cross-View Training
2018
Unsupervised representation learning algorithms such as word2vec and ELMo improve the accuracy of many supervised NLP models, mainly because they can take advantage of large amounts of unlabeled text. However, the supervised models only learn …
-
Effective Approaches to Attention-based Neural Machine Translation
2015 · arXiv (Cornell University)
An attentional mechanism has lately been used to improve neural machine translation (NMT) by selectively focusing on parts of the source sentence during translation. However, there has been little work exploring useful architectures for attention-based …
-
A Hierarchical Neural Autoencoder for Paragraphs and Documents
2015 · arXiv (Cornell University)
Jiwei Li, Thang Luong, Dan Jurafsky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
-
When Are Tree Structures Necessary for Deep Learning of Representations?
2015 · arXiv (Cornell University)
Recursive neural models, which use syntactic parse trees to recursively generate representations bottom-up, are a popular architecture. But there have not been rigorous evaluations showing for exactly which tasks this syntax-based method is appropriate. In …
-
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
2020 · arXiv (Cornell University)
Masked language modeling (MLM) pre-training methods such as BERT corrupt the input by replacing some tokens with [MASK] and then train a model to reconstruct the original tokens. While they produce good results when transferred …
-
Towards a Human-like Open-Domain Chatbot
2020 · arXiv (Cornell University)
We present Meena, a multi-turn open-domain chatbot trained end-to-end on data mined and filtered from public domain social media conversations. This 2.6B parameter neural network is simply trained to minimize perplexity of the next token. …