Researcher profile

Minh-Thang Luong

11 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Multi-task Sequence to Sequence Learning

    2015 · arXiv (Cornell University)

    Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. …

  2. Achieving Open Vocabulary Neural Machine Translation with Hybrid Word-Character Models

    2016

    Nearly all previous work on neural machine translation (NMT) has used quite restricted vocabularies, perhaps with a subsequent method to patch in unknown words. This paper presents a novel wordcharacter solution to achieving open vocabulary …

  3. Massive Exploration of Neural Machine Translation Architectures

    2017 · arXiv (Cornell University)

    Neural Machine Translation (NMT) has shown remarkable progress over the past few years, with production systems now being deployed to end-users. As the field is moving rapidly, it has become unclear which elements of NMT …

  4. Online and Linear-Time Attention by Enforcing Monotonic Alignments

    2017 · arXiv (Cornell University)

    Recurrent neural network models with an attention mechanism have proven to be extremely effective on a wide variety of sequence-to-sequence problems. However, the fact that soft attention mechanisms perform a pass over the entire input …

  5. QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension

    2018 · arXiv (Cornell University)

    Current end-to-end machine reading and question answering (Q\&A) models are primarily based on recurrent neural networks (RNNs) with attention. Despite their success, these models are often slow for both training and inference due to the …

  6. Semi-Supervised Sequence Modeling with Cross-View Training

    2018

    Unsupervised representation learning algorithms such as word2vec and ELMo improve the accuracy of many supervised NLP models, mainly because they can take advantage of large amounts of unlabeled text. However, the supervised models only learn …

  7. Effective Approaches to Attention-based Neural Machine Translation

    2015 · arXiv (Cornell University)

    An attentional mechanism has lately been used to improve neural machine translation (NMT) by selectively focusing on parts of the source sentence during translation. However, there has been little work exploring useful architectures for attention-based …

  8. A Hierarchical Neural Autoencoder for Paragraphs and Documents

    2015 · arXiv (Cornell University)

    Jiwei Li, Thang Luong, Dan Jurafsky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  9. When Are Tree Structures Necessary for Deep Learning of Representations?

    2015 · arXiv (Cornell University)

    Recursive neural models, which use syntactic parse trees to recursively generate representations bottom-up, are a popular architecture. But there have not been rigorous evaluations showing for exactly which tasks this syntax-based method is appropriate. In …

  10. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

    2020 · arXiv (Cornell University)

    Masked language modeling (MLM) pre-training methods such as BERT corrupt the input by replacing some tokens with [MASK] and then train a model to reconstruct the original tokens. While they produce good results when transferred …

  11. Towards a Human-like Open-Domain Chatbot

    2020 · arXiv (Cornell University)

    We present Meena, a multi-turn open-domain chatbot trained end-to-end on data mined and filtered from public domain social media conversations. This 2.6B parameter neural network is simply trained to minimize perplexity of the next token. …