Researcher profile

Richard Socher

26 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Deep Learning for Sentiment Analysis - Invited Talk

    2016

    Richard Socher is the CEO and founder of MetaMind, a startup that seeks to improve artificial intelligence and make it widely accessible. He obtained his PhD from Stanford working on deep learning with Chris Manning …

  2. A Deep Reinforced Model for Abstractive Summarization

    2017 · arXiv (Cornell University)

    Attentional, RNN-based encoder-decoder models for abstractive summarization have achieved good performance on short input and output sequences. For longer documents and summaries however these models often include repetitive and incoherent phrases. We introduce a neural …

  3. Efficient and Robust Question Answering from Minimal Context over Documents

    2018 · ArXiv.org

    Neural models for question answering (QA) over documents have achieved significant performance improvements. Although effective, these models do not scale to large corpora due to their complex modeling of interactions between the document and the …

  4. Coarse-grain Fine-grain Coattention Network for Multi-evidence Question Answering

    2019 · International Conference on Learning Representations

    End-to-end neural models have made significant progress in question answering, however recent studies show that these models implicitly assume that the answer and evidence appear close together in a single document. In this work, we …

  5. DivideMix: Learning with Noisy Labels as Semi-supervised Learning

    2020 · arXiv (Cornell University)

    Deep neural networks are known to be annotation-hungry. Numerous efforts have been devoted to reducing the annotation cost when learning with deep networks. Two prominent directions include learning with noisy labels and semi-supervised learning by …

  6. GeDi: Generative Discriminator Guided Sequence Generation

    2020 · arXiv (Cornell University)

    While large-scale language models (LMs) are able to imitate the distribution of natural language well enough to generate realistic text, it is difficult to control which regions of the distribution they generate. This is especially …

  7. Prostate cancer therapy personalization via multi-modal deep learning on randomized phase III clinical trials

    2022 · Research Square

    <title>Abstract</title> Prostate cancer is the most frequent cancer in men and a leading cause of cancer death. Determining a patient’s optimal therapy is a challenge, where oncologists must select a therapy with the highest likelihood …

  8. Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks

    2015

    Kai Sheng Tai, Richard Socher, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  9. Ask Me Anything: Dynamic Memory Networks for Natural Language Processing

    2015 · arXiv (Cornell University)

    Most tasks in natural language processing can be cast into question answering (QA) problems over language input. We introduce the dynamic memory network (DMN), a neural network architecture which processes input sequences and questions, forms …

  10. Pointer Sentinel Mixture Models

    2016 · arXiv (Cornell University)

    Recent neural network sequence models with softmax classifiers have achieved their best language modeling performance only with very large hidden states and large vocabularies. Even then they struggle to predict rare or unseen words even …

  11. Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling

    2016 · arXiv (Cornell University)

    Recurrent neural networks have been very successful at predicting sequences of words in tasks such as language modeling. However, all such models are based on the conventional classification framework, where the model is trained against …

  12. Dynamic Coattention Networks For Question Answering

    2016 · arXiv (Cornell University)

    Several deep learning models have been proposed for question answering. However, due to their single-pass nature, they have no way to recover from local maxima corresponding to incorrect answers. To address this problem, we introduce …

  13. A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks

    2017

    Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks. Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in …

  14. Learned in Translation: Contextualized Word Vectors

    2017 · arXiv (Cornell University)

    Computer vision has benefited from initializing multiple deep layers with weights pretrained on large supervised training sets like ImageNet. Natural language processing (NLP) typically sees initialization of only the lowest layer of deep models with …

  15. DCN+: Mixed Objective and Deep Residual Coattention for Question Answering

    2017 · arXiv (Cornell University)

    Traditional models for question answering optimize using cross entropy loss, which encourages exact answers at the cost of penalizing nearby or overlapping answers that are sometimes equally accurate. We propose a mixed objective that combines …

  16. Non-Autoregressive Neural Machine Translation

    2017 · arXiv (Cornell University)

    Existing approaches to neural machine translation condition each output word on previously generated outputs. We introduce a model that avoids this autoregressive property and produces its outputs in parallel, allowing an order of magnitude lower …

  17. An Analysis of Neural Language Modeling at Multiple Scales

    2018 · arXiv (Cornell University)

    Many of the leading approaches in language modeling introduce novel, complex and specialized architectures. We take existing state-of-the-art word level language models based on LSTMs and QRNNs and extend them to both larger vocabularies as …

  18. The Natural Language Decathlon: Multitask Learning as Question Answering

    2018 · arXiv (Cornell University)

    Deep learning has improved performance on many natural language processing (NLP) tasks individually. However, general NLP models cannot emerge within a paradigm that focuses on the particularities of a single metric, dataset, and task. We …

  19. XLDA: Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering

    2019 · arXiv (Cornell University)

    While natural language processing systems often focus on a single language, multilingual transfer learning has the potential to improve performance, especially for low-resource languages. We introduce XLDA, cross-lingual data augmentation, a method that replaces a …

  20. A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks

    2016 · arXiv (Cornell University)

    Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks. Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in …

  21. Quasi-Recurrent Neural Networks

    2016 · arXiv (Cornell University)

    Recurrent neural networks are a powerful tool for modeling sequential data, but the dependence of each timestep's computation on the previous timestep's output limits parallelism and makes RNNs unwieldy for very long sequences. We introduce …

  22. CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases

    2019

    Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao …

  23. CTRL: A Conditional Transformer Language Model for Controllable Generation

    2019 · arXiv (Cornell University)

    Large-scale language models show promising text generation capabilities, but users cannot easily control particular aspects of the generated text. We release CTRL, a 1.63 billion-parameter conditional transformer language model, trained to condition on control codes …

  24. BERT is Not an Interlingua and the Bias of Tokenization

    2019

    Multilingual transfer learning can benefit both high-and low-resource languages, but the source of these improvements is not well understood. Cananical Correlation Analysis (CCA) of the internal representations of a pretrained, multilingual BERT model reveals that …

  25. Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering

    2019 · arXiv (Cornell University)

    Answering questions that require multi-hop reasoning at web-scale necessitates retrieving multiple evidence documents, one of which often has little lexical or semantic relationship to the question. This paper introduces a new graph-based recurrent retrieval approach …

  26. GeDi: Generative Discriminator Guided Sequence Generation

    2021

    Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, Nazneen Fatema Rajani. Findings of the Association for Computational Linguistics: EMNLP 2021. 2021.