ملف الباحث

Ruslan Salakhutdinov

15 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations

    2016 · arXiv (Cornell University)

    We investigate the parameter-space geometry of recurrent neural networks (RNNs), and develop an adaptation of path-SGD optimization method, attuned to this geometry, that can learn plain RNNs with ReLU activations. On several datasets that require …

  2. Normalized Gradient with Adaptive Stepsize Method for Deep Neural Network Training.

    2017 · arXiv (Cornell University)

    In this paper, we propose a generic and simple algorithmic framework for first order optimization. The framework essentially contains two consecutive steps in each iteration: 1) computing and normalizing the mini-batch stochastic gradient; 2) selecting …

  3. Strong and Simple Baselines for Multimodal Utterance Embeddings

    2019

    Paul Pu Liang, Yao Chong Lim, Yao-Hung Hubert Tsai, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long …

  4. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

    2018 · arXiv (Cornell University)

    Existing question answering (QA) datasets fail to train QA systems to perform complex reasoning and provide explanations for answers. We introduce HotpotQA, a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) …

  5. Skip-Thought Vectors

    2015 · arXiv (Cornell University)

    We describe an approach for unsupervised learning of a generic, distributed sentence encoder. Using the continuity of text from books, we train an encoder-decoder model that tries to reconstruct the surrounding sentences of an encoded …

  6. Multi-Task Cross-Lingual Sequence Tagging from Scratch

    2016 · arXiv (Cornell University)

    We present a deep hierarchical recurrent neural network for sequence tagging. Given a sequence of words, our model employs deep gated recurrent units on both character and word levels to encode morphology and context information, …

  7. Transfer Learning for Sequence Tagging with Hierarchical Recurrent Networks

    2017 · arXiv (Cornell University)

    Recent papers have shown that neural networks obtain state-of-the-art performance on several different sequence tagging tasks. One appealing property of such systems is their generality, as excellent performance can be achieved with a unified architecture …

  8. Gated-Attention Architectures for Task-Oriented Language Grounding

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    To perform tasks specified by natural language instructions, autonomous agents need to extract semantically meaningful representations of language and map it to visual elements and actions in the environment. This problem is called task-oriented language …

  9. Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

    2017 · arXiv (Cornell University)

    We formulate language modeling as a matrix factorization problem, and show that the expressiveness of Softmax-based models (including the majority of neural language models) is limited by a Softmax bottleneck. Given that natural language is …

  10. Style Transfer Through Back-Translation

    2018

    Style transfer is the task of rephrasing the text to contain specific stylistic properties without changing the intent or affect within the context. This paper introduces a new method for automatic style transfer. We first …

  11. Open Domain Question Answering Using Early Fusion of Knowledge Bases and Text

    2018

    Open Domain Question Answering (QA) is evolving from complex pipelined systems to end-to-end deep neural networks. Specialized neural models have been developed for extracting answers from either text alone or Knowledge Bases (KBs) alone. In …

  12. Bootstrapping a Data-Set and Model for Question-Answering in Portuguese (Short Paper)

    2019 · arXiv (Cornell University)

    Question answering systems are mainly concerned with fulfilling an information query written in natural language, given a collection of documents with relevant information. They are key elements in many popular application systems as personal assistants, …

  13. XLNet: Generalized Autoregressive Pretraining for Language Understanding

    2019 · arXiv (Cornell University)

    With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency …

  14. Gated-Attention Architectures for Task-Oriented Language Grounding

    2017 · arXiv (Cornell University)

    To perform tasks specified by natural language instructions, autonomous agents need to extract semantically meaningful representations of language and map it to visual elements and actions in the environment. This problem is called task-oriented language …

  15. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context

    2019

    Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed …