Researcher profile

Yoshua Bengio

25 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Sparse Attentive Backtracking: Long-Range Credit Assignment in Recurrent Networks

    2017 · arXiv (Cornell University)

    A major drawback of backpropagation through time (BPTT) is the difficulty of learning long-term dependencies, coming from having to propagate credit information backwards through every single step of the forward computation. This makes BPTT both …

  2. Learning to Understand Phrases by Embedding the Dictionary

    2015 · arXiv (Cornell University)

    Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …

  3. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

    2018 · arXiv (Cornell University)

    Existing question answering (QA) datasets fail to train QA systems to perform complex reasoning and provide explanations for answers. We introduce HotpotQA, a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) …

  4. A Structured Self-Attentive Sentence Embedding.

    2017 · International Conference on Learning Representations

    This paper proposes a new model for extracting an interpretable sentence embedding by introducing self-attention. Instead of using a vector, we use a 2-D matrix to represent the embedding, with each row of the matrix …

  5. Systematic Evaluation of Causal Discovery in Visual Model Based Reinforcement Learning

    2021 · arXiv (Cornell University)

    Inducing causal relationships from observations is a classic problem in machine learning. Most work in causality starts from the premise that the causal variables themselves are observed. However, for AI agents such as robots trying …

  6. The Causal-Neural Connection: Expressiveness, Learnability, and\n Inference

    2021 · arXiv (Cornell University)

    One of the central elements of any causal inference is an object called\nstructural causal model (SCM), which represents a collection of mechanisms and\nexogenous sources of random variation of the system under investigation (Pearl,\n2000). An important …

  7. BatchGFN: Generative Flow Networks for Batch Active Learning

    2023 · arXiv (Cornell University)

    We introduce BatchGFN -- a novel approach for pool-based active learning that uses generative flow networks to sample sets of data points proportional to a batch reward. With an appropriate reward function to quantify the …

  8. Leveraging Diffusion Disentangled Representations to Mitigate Shortcuts in Underspecified Visual Tasks

    2023 · arXiv (Cornell University)

    Spurious correlations in the data, where multiple cues are predictive of the target labels, often lead to shortcut learning phenomena, where a model may rely on erroneous, easy-to-learn, cues while ignoring reliable ones. In this …

  9. Learning to Understand Phrases by Embedding the Dictionary

    2016 · Transactions of the Association for Computational Linguistics

    Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …

  10. Gated Feedback Recurrent Neural Networks

    2015 · arXiv (Cornell University)

    In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper …

  11. On Using Monolingual Corpora in Neural Machine Translation

    2015 · HAL (Le Centre pour la Communication Scientifique Directe)

    Recent work on end-to-end neural network-based architectures for machine translation has shown promising results for En-Fr and En-De translation. Arguably, one of the major factors behind this success has been the availability of high quality …

  12. A Hierarchical Recurrent Encoder-Decoder for Generative Context-Aware Query Suggestion

    2015

    Users may strive to formulate an adequate textual query for their information need. Search engines assist the users by presenting query suggestions. To preserve the original search intent, suggestions should be context-aware and account for …

  13. On Using Very Large Target Vocabulary for Neural Machine Translation

    2015

    Sébastien Jean, Kyunghyun Cho, Roland Memisevic, Yoshua Bengio. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  14. End-to-end attention-based large vocabulary speech recognition

    2016

    Many state-of-the-art Large Vocabulary Continuous Speech Recognition (LVCSR) Systems are hybrids of neural networks and Hidden Markov Models (HMMs). Recently, more direct end-to-end methods have been investigated, in which neural architectures were trained to model …

  15. Multi-Way, Multilingual Neural Machine Translation with a Shared Attention Mechanism

    2016

    We propose multi-way, multilingual neural machine translation. The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages. …

  16. Montreal Neural Machine Translation Systems for WMT’15

    2015

    Neural machine translation (NMT) systems have recently achieved results comparable to the state of the art on a few translation tasks, including EnglishFrench and EnglishGerman. The main purpose of the Montreal Institute for Learning Algorithms …

  17. A Character-level Decoder without Explicit Segmentation for Neural Machine Translation

    2016

    The existing machine translation systems, whether phrase-based or neural, have relied almost exclusively on word-level modelling with explicit segmentation. In this paper, we ask a fundamental question: can neural machine translation generate a character sequence …

  18. Pointing the Unknown Words

    2016 · arXiv (Cornell University)

    The problem of rare and unknown words is an important issue that can potentially influence the performance of many NLP systems, including both the traditional count-based and the deep learning models. We propose a novel …

  19. Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations

    2016 · PolyPublie (École Polytechnique de Montréal)

    We propose zoneout, a novel method for regularizing RNNs. At each timestep, zoneout stochastically forces some hidden units to maintain their previous values. Like dropout, zoneout uses random noise to train a pseudo-ensemble, improving generalization. …

  20. Multiresolution Recurrent Neural Networks: An Application to Dialogue Response Generation

    2017 · Proceedings of the AAAI Conference on Artificial Intelligence

    We introduce a new class of models called multiresolution recurrent neural networks, which explicitly model natural language generation at multiple levels of abstraction. The models extend the sequence-to-sequence framework to generate two parallel stochastic processes: …

  21. An Actor-Critic Algorithm for Sequence Prediction

    2016 · arXiv (Cornell University)

    We present an approach to training neural networks to generate sequences using actor-critic methods from reinforcement learning (RL). Current log-likelihood training methods are limited by the discrepancy between their training and testing modes, as models …

  22. Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning

    2018 · PolyPublie (École Polytechnique de Montréal)

    A lot of the recent success in natural language processing (NLP) has been driven by distributed vector representations of words trained on large amounts of text in an unsupervised manner. These representations are typically used …

  23. Attention-Based Models for Speech Recognition

    2015 · arXiv (Cornell University)

    Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend …

  24. Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models

    2016 · Proceedings of the AAAI Conference on Artificial Intelligence

    We investigate the task of building open domain, conversational dialogue systems based on large dialogue corpora using generative models. Generative models produce system responses that are autonomously generated word-by-word, opening up the possibility for realistic, …

  25. Deep Graph Infomax

    2018 · Apollo (University of Cambridge)

    We present Deep Graph Infomax (DGI), a general approach for learning node representations within graph-structured data in an unsupervised manner. DGI relies on maximizing mutual information between patch representations and corresponding high-level summaries of graphs---both …