Yoshua Bengio
25 papers in the PaperMetrix corpus
Papers by this author
-
Sparse Attentive Backtracking: Long-Range Credit Assignment in Recurrent Networks
2017 · arXiv (Cornell University)
A major drawback of backpropagation through time (BPTT) is the difficulty of learning long-term dependencies, coming from having to propagate credit information backwards through every single step of the forward computation. This makes BPTT both …
-
Learning to Understand Phrases by Embedding the Dictionary
2015 · arXiv (Cornell University)
Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …
-
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
2018 · arXiv (Cornell University)
Existing question answering (QA) datasets fail to train QA systems to perform complex reasoning and provide explanations for answers. We introduce HotpotQA, a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) …
-
A Structured Self-Attentive Sentence Embedding.
2017 · International Conference on Learning Representations
This paper proposes a new model for extracting an interpretable sentence embedding by introducing self-attention. Instead of using a vector, we use a 2-D matrix to represent the embedding, with each row of the matrix …
-
Systematic Evaluation of Causal Discovery in Visual Model Based Reinforcement Learning
2021 · arXiv (Cornell University)
Inducing causal relationships from observations is a classic problem in machine learning. Most work in causality starts from the premise that the causal variables themselves are observed. However, for AI agents such as robots trying …
-
The Causal-Neural Connection: Expressiveness, Learnability, and\n Inference
2021 · arXiv (Cornell University)
One of the central elements of any causal inference is an object called\nstructural causal model (SCM), which represents a collection of mechanisms and\nexogenous sources of random variation of the system under investigation (Pearl,\n2000). An important …
-
BatchGFN: Generative Flow Networks for Batch Active Learning
2023 · arXiv (Cornell University)
We introduce BatchGFN -- a novel approach for pool-based active learning that uses generative flow networks to sample sets of data points proportional to a batch reward. With an appropriate reward function to quantify the …
-
Leveraging Diffusion Disentangled Representations to Mitigate Shortcuts in Underspecified Visual Tasks
2023 · arXiv (Cornell University)
Spurious correlations in the data, where multiple cues are predictive of the target labels, often lead to shortcut learning phenomena, where a model may rely on erroneous, easy-to-learn, cues while ignoring reliable ones. In this …
-
Learning to Understand Phrases by Embedding the Dictionary
2016 · Transactions of the Association for Computational Linguistics
Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …
-
Gated Feedback Recurrent Neural Networks
2015 · arXiv (Cornell University)
In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper …
-
On Using Monolingual Corpora in Neural Machine Translation
2015 · HAL (Le Centre pour la Communication Scientifique Directe)
Recent work on end-to-end neural network-based architectures for machine translation has shown promising results for En-Fr and En-De translation. Arguably, one of the major factors behind this success has been the availability of high quality …
-
A Hierarchical Recurrent Encoder-Decoder for Generative Context-Aware Query Suggestion
2015
Users may strive to formulate an adequate textual query for their information need. Search engines assist the users by presenting query suggestions. To preserve the original search intent, suggestions should be context-aware and account for …
-
On Using Very Large Target Vocabulary for Neural Machine Translation
2015
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, Yoshua Bengio. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
-
End-to-end attention-based large vocabulary speech recognition
2016
Many state-of-the-art Large Vocabulary Continuous Speech Recognition (LVCSR) Systems are hybrids of neural networks and Hidden Markov Models (HMMs). Recently, more direct end-to-end methods have been investigated, in which neural architectures were trained to model …
-
Multi-Way, Multilingual Neural Machine Translation with a Shared Attention Mechanism
2016
We propose multi-way, multilingual neural machine translation. The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages. …
-
Montreal Neural Machine Translation Systems for WMT’15
2015
Neural machine translation (NMT) systems have recently achieved results comparable to the state of the art on a few translation tasks, including EnglishFrench and EnglishGerman. The main purpose of the Montreal Institute for Learning Algorithms …
-
A Character-level Decoder without Explicit Segmentation for Neural Machine Translation
2016
The existing machine translation systems, whether phrase-based or neural, have relied almost exclusively on word-level modelling with explicit segmentation. In this paper, we ask a fundamental question: can neural machine translation generate a character sequence …
-
Pointing the Unknown Words
2016 · arXiv (Cornell University)
The problem of rare and unknown words is an important issue that can potentially influence the performance of many NLP systems, including both the traditional count-based and the deep learning models. We propose a novel …
-
Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
2016 · PolyPublie (École Polytechnique de Montréal)
We propose zoneout, a novel method for regularizing RNNs. At each timestep, zoneout stochastically forces some hidden units to maintain their previous values. Like dropout, zoneout uses random noise to train a pseudo-ensemble, improving generalization. …
-
Multiresolution Recurrent Neural Networks: An Application to Dialogue Response Generation
2017 · Proceedings of the AAAI Conference on Artificial Intelligence
We introduce a new class of models called multiresolution recurrent neural networks, which explicitly model natural language generation at multiple levels of abstraction. The models extend the sequence-to-sequence framework to generate two parallel stochastic processes: …
-
An Actor-Critic Algorithm for Sequence Prediction
2016 · arXiv (Cornell University)
We present an approach to training neural networks to generate sequences using actor-critic methods from reinforcement learning (RL). Current log-likelihood training methods are limited by the discrepancy between their training and testing modes, as models …
-
Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
2018 · PolyPublie (École Polytechnique de Montréal)
A lot of the recent success in natural language processing (NLP) has been driven by distributed vector representations of words trained on large amounts of text in an unsupervised manner. These representations are typically used …
-
Attention-Based Models for Speech Recognition
2015 · arXiv (Cornell University)
Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend …
-
Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models
2016 · Proceedings of the AAAI Conference on Artificial Intelligence
We investigate the task of building open domain, conversational dialogue systems based on large dialogue corpora using generative models. Generative models produce system responses that are autonomously generated word-by-word, opening up the possibility for realistic, …
-
Deep Graph Infomax
2018 · Apollo (University of Cambridge)
We present Deep Graph Infomax (DGI), a general approach for learning node representations within graph-structured data in an unsupervised manner. DGI relies on maximizing mutual information between patch representations and corresponding high-level summaries of graphs---both …