Researcher profile

Jason Weston

22 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Dialogue Natural Language Inference

    2019

    Consistency is a long standing issue faced by dialogue models. In this paper, we frame the consistency of dialogue agents as natural language inference (NLI) and create a new natural language inference dataset called Dialogue …

  2. I love your chain mail! Making knights smile in a fantasy game world:\n Open-domain goal-oriented dialogue agents

    2020 · arXiv (Cornell University)

    Dialogue research tends to distinguish between chit-chat and goal-oriented\ntasks. While the former is arguably more naturalistic and has a wider use of\nlanguage, the latter has clearer metrics and a straightforward learning signal.\nHumans effortlessly combine the …

  3. Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

    2015 · arXiv (Cornell University)

    One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue for …

  4. End-To-End Memory Networks

    2015 · arXiv (Cornell University)

    We introduce a neural network with a recurrent attention model over a possibly large external memory. The architecture is a form of Memory Network (Weston et al., 2015) but unlike the model in that work, …

  5. A Neural Attention Model for Abstractive Sentence Summarization

    2015

    Summarization based on text extraction is inherently limited, but generation-style abstractive methods have proven challenging to build. In this work, we propose a fully data-driven approach to abstractive sentence summarization. Our method utilizes a local …

  6. Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution

    2015

    Sam Wiseman, Alexander M. Rush, Stuart Shieber, Jason Weston. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). …

  7. A Neural Attention Model for Sentence Summarization

    2015

    Summarization based on text extraction is inherently limited, but generation-style ab-stractive methods have proven challeng-ing to build. In this work, we propose a fully data-driven approach to abstrac-tive sentence summarization. Our method utilizes a local …

  8. Key-Value Memory Networks for Directly Reading Documents

    2016

    Directly reading documents and being able to answer questions from them is an unsolved challenge. To avoid its inherent difficulty, question answering (QA) has been directed towards using Knowledge Bases (KBs) instead, which has proven …

  9. Tracking the World State with Recurrent Entity Networks

    2016 · arXiv (Cornell University)

    We introduce a new model, the Recurrent Entity Network (EntNet). It is equipped with a dynamic long-term memory which allows it to maintain and update a representation of the state of the world as it …

  10. Reading Wikipedia to Answer Open-Domain Questions

    2017 · arXiv (Cornell University)

    This paper proposes to tackle open- domain question answering using Wikipedia as the unique knowledge source: the answer to any factoid question is a text span in a Wikipedia article. This task of machine reading …

  11. Personalizing Dialogue Agents: I have a dog, do you have pets too?

    2018 · arXiv (Cornell University)

    Chit-chat models are known to have several problems: they lack specificity, do not display a consistent personality and are often not very captivating. In this work we present the task of making chit-chat more engaging …

  12. Retrieve and Refine: Improved Sequence Generation Models For Dialogue

    2018

    Sequence generation models for dialogue are known to have several problems: they tend to produce short, generic sentences that are uninformative and unengaging. Retrieval models on the other hand can surface interesting responses, but are …

  13. Wizard of Wikipedia: Knowledge-Powered Conversational agents

    2018 · arXiv (Cornell University)

    In open-domain dialogue intelligent agents should exhibit the use of knowledge, however there are few convincing demonstrations of this to date. The most popular sequence to sequence models typically "generate and hope" generic utterances that …

  14. What makes a good conversation? How controllable attributes affect human judgments

    2019

    Abigail See, Stephen Roller, Douwe Kiela, Jason Weston. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.

  15. ELI5: Long Form Question Answering

    2019

    We introduce the first large-scale corpus for long-form question answering, a task requiring elaborate and in-depth answers to openended questions. The dataset comprises 270K threads from the Reddit forum "Explain Like I'm Five" (ELI5) where …

  16. Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

    2016 · International Conference on Learning Representations

    Abstract: One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue …

  17. The Goldilocks Principle: Reading Children's Books with Explicit Memory Representations

    2016 · arXiv (Cornell University)

    Abstract: We introduce a new test of how well language models capture meaning in children's books. Unlike standard language modelling benchmarks, it distinguishes the task of predicting syntactic function words from that of predicting lower-frequency …

  18. Neural Text Generation with Unlikelihood Training

    2019 · arXiv (Cornell University)

    Neural text generation is a key tool in natural language applications, but it is well known there are major problems at its core. In particular, standard likelihood training and decoding leads to dull and repetitive …

  19. Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring

    2020 · International Conference on Learning Representations

    The use of deep pre-trained transformers has led to remarkable progress in a number of applications (Devlin et al., 2018). For tasks that make pairwise comparisons between sequences, matching a given input with a corresponding …

  20. Recipes for Building an Open-Domain Chatbot

    2021

    Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston. Proceedings of the 16th Conference of the European Chapter of the Association …

  21. Retrieval Augmentation Reduces Hallucination in Conversation

    2021

    Despite showing increasingly human-like conversational abilities, state-of-the-art dialogue models often suffer from factual incorrectness and hallucination of knowledge (Roller et al., 2020). In this work we explore the use of neural-retrieval-in-the-loop architectures - recently shown …

  22. Large-scale Simple Question Answering with Memory Networks

    2015 · arXiv (Cornell University)

    Training large-scale question answering systems is complicated because training sources usually cover a small portion of the range of possible questions. This paper studies the impact of multitask and transfer learning for simple question answering; …