Quoc V. Le
35 papers in the PaperMetrix corpus
Papers by this author
-
Effective Domain Mixing for Neural Machine Translation
2017
Neural Machine Translation (NMT) models are often trained on heterogeneous mixtures of domains, from news to parliamentary proceedings, each with unique distributions and language. In this work we show that training NMT systems on naively …
-
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
2017 · arXiv (Cornell University)
We consider two questions at the heart of machine learning; how can we predict if a minimum will generalize to the test set, and why does stochastic gradient descent find minima that generalize well? Our …
-
Evolving modular neural sequence architectures with genetic programming
2018 · Proceedings of the Genetic and Evolutionary Computation Conference Companion
Automated architecture search has demonstrated significant success for image data, where reinforcement learning and evolution approaches now outperform the best human designed networks ([12], [8]). These successes have not transferred over to models dealing with …
-
Memory Augmented Policy Optimization for Program Synthesis with Generalization
2018 · arXiv (Cornell University)
This paper presents Memory Augmented Policy Optimization (MAPO): a novel policy optimization formulation that incorporates a memory buffer of promising trajectories to reduce the variance of policy gradient estimates for deterministic environments with discrete actions. …
-
Towards a Reliable and Context-Based System Architecture for Autonomous Vehicles
2020
Full vehicle autonomy excludes a takeover by passengers in case a safety-critical application fails. Therefore, the system responsible for operating the autonomous vehicle has to detect and handle failures autonomously. Moreover, this system has to …
-
Evolving Reinforcement Learning Algorithms
2021 · arXiv (Cornell University)
We propose a method for meta-learning reinforcement learning algorithms by searching over the space of computational graphs which compute the loss function for a value-based model-free RL agent to optimize. The learned algorithms are domain-agnostic …
-
Symbol tuning improves in-context learning in language models
2023 · arXiv (Cornell University)
We present symbol tuning - finetuning language models on in-context input-label pairs where natural language labels (e.g., "positive/negative sentiment") are replaced with arbitrary symbols (e.g., "foo/bar"). Symbol tuning leverages the intuition that when a model …
-
A Neural Conversational Model
2015 · arXiv (Cornell University)
Conversational modeling is an important task in natural language understanding and machine intelligence. Although previous approaches exist, they are often restricted to specific domains (e.g., booking an airline ticket) and require hand-crafted rules. In this …
-
A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
2015 · arXiv (Cornell University)
Learning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients. To overcome this difficulty, researchers have developed sophisticated optimization techniques and network architectures. In this paper, we propose a simpler …
-
Addressing the Rare Word Problem in Neural Machine Translation
2015
Thang Luong, Ilya Sutskever, Quoc Le, Oriol Vinyals, Wojciech Zaremba. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long …
-
Semi-supervised Sequence Learning
2015 · arXiv (Cornell University)
We present two approaches that use unlabeled data to improve sequence learning with recurrent networks. The first approach is to predict what comes next in a sequence, which is a conventional language model in natural …
-
Multi-task Sequence to Sequence Learning
2015 · arXiv (Cornell University)
Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. …
-
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
2016
We present Listen, Attend and Spell (LAS), a neural speech recognizer that transcribes speech utterances directly to characters without pronunciation models, HMMs or other components of traditional speech recognizers. In LAS, the neural network architecture …
-
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
2016 · arXiv (Cornell University)
Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive …
-
Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak Supervision
2017
Harnessing the statistical power of neural networks to perform language understanding and symbolic reasoning is difficult, when it requires executing efficient discrete operations against a large knowledge-base. In this work, we introduce a Neural Symbolic …
-
Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation
2017 · Transactions of the Association for Computational Linguistics
We propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no changes to the model architecture from a standard NMT system but instead …
-
Unsupervised Pretraining for Sequence to Sequence Learning
2017
This work presents a general unsupervised learning method to improve the accuracy of sequence to sequence (seq2seq) models. In our method, the weights of the encoder and decoder of a seq2seq model are initialized with …
-
Massive Exploration of Neural Machine Translation Architectures
2017 · arXiv (Cornell University)
Neural Machine Translation (NMT) has shown remarkable progress over the past few years, with production systems now being deployed to end-users. As the field is moving rapidly, it has become unclear which elements of NMT …
-
QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
2018 · arXiv (Cornell University)
Current end-to-end machine reading and question answering (Q\&A) models are primarily based on recurrent neural networks (RNNs) with attention. Despite their success, these models are often slow for both training and inference due to the …
-
A Simple Method for Commonsense Reasoning
2018 · arXiv (Cornell University)
Commonsense reasoning is a long-standing challenge for deep learning. For example, it is difficult to use neural networks to tackle the Winograd Schema dataset (Levesque et al., 2011). In this paper, we present a simple …
-
Semi-Supervised Sequence Modeling with Cross-View Training
2018
Unsupervised representation learning algorithms such as word2vec and ELMo improve the accuracy of many supervised NLP models, mainly because they can take advantage of large amounts of unlabeled text. However, the supervised models only learn …
-
Bootstrapping a Data-Set and Model for Question-Answering in Portuguese (Short Paper)
2019 · arXiv (Cornell University)
Question answering systems are mainly concerned with fulfilling an information query written in natural language, given a collection of documents with relevant information. They are key elements in many popular application systems as personal assistants, …
-
Natural Questions: A Benchmark for Question Answering Research
2019 · Transactions of the Association for Computational Linguistics
We present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia …
-
Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak Supervision
2016 · arXiv (Cornell University)
Harnessing the statistical power of neural networks to perform language understanding and symbolic reasoning is difficult, when it requires executing efficient discrete operations against a large knowledge-base. In this work, we introduce a Neural Symbolic …
-
XLNet: Generalized Autoregressive Pretraining for Language Understanding
2019 · arXiv (Cornell University)
With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency …
-
Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation
2016 · arXiv (Cornell University)
We propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no change in the model architecture from our base system but instead introduces …
-
Neural Architecture Search with Reinforcement Learning
2016 · arXiv (Cornell University)
Neural networks are powerful and flexible models that work well for many difficult learning tasks in image, speech and natural language understanding. Despite their success, neural networks are still hard to design. In this paper, …
-
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
2019
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed …
-
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
2020 · arXiv (Cornell University)
Masked language modeling (MLM) pre-training methods such as BERT corrupt the input by replacing some tokens with [MASK] and then train a model to reconstruct the original tokens. While they produce good results when transferred …
-
Towards a Human-like Open-Domain Chatbot
2020 · arXiv (Cornell University)
We present Meena, a multi-turn open-domain chatbot trained end-to-end on data mined and filtered from public domain social media conversations. This 2.6B parameter neural network is simply trained to minimize perplexity of the next token. …
-
BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
2022 · arXiv (Cornell University)
AbstractThere is a failure mode in large language models that we do not have a good name for, and thatwe therefore tend not to treat seriously enough. It is not hallucination — the model is …
-
Self-Consistency Improves Chain of Thought Reasoning in Language Models
2022 · arXiv (Cornell University)
Chain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks. In this paper, we propose a new decoding strategy, self-consistency, to replace the naive greedy decoding used in chain-of-thought …
-
LaMDA: Language Models for Dialog Applications
2022 · arXiv (Cornell University)
We present LaMDA: Language Models for Dialog Applications. LaMDA is a family of Transformer-based neural language models specialized for dialog, which have up to 137B parameters and are pre-trained on 1.56T words of public dialog …
-
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
2022 · arXiv (Cornell University)
Chain-of-thought prompting has demonstrated remarkable performance on various natural language reasoning tasks. However, it tends to perform poorly on tasks which requires solving problems harder than the exemplars shown in the prompts. To overcome this …
-
Scaling Instruction-Finetuned Language Models
2022 · arXiv (Cornell University)
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on …