ملف الباحث

Alexander M. Rush

25 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. On Adversarial Removal of Hypothesis-only Bias in Natural Language Inference

    2019

    Popular Natural Language Inference (NLI) datasets have been shown to be tainted by hypothesis-only biases. Adversarial learning may help models ignore sensitive biases and spurious correlations in data. We evaluate whether adversarial learning can be …

  2. Sequence-Level Mixed Sample Data Augmentation

    2020

    Despite their empirical success, neural networks still have difficulty capturing compositional aspects of natural language. This work proposes a simple data augmentation approach to encourage compositional behavior in neural models for sequence-to-sequence problems. Our approach, …

  3. GRIT: Generative Role-filler Transformers for Document-level Event Entity Extraction

    2021

    We revisit the classic problem of documentlevel role-filler entity extraction (REE) for template filling. We argue that sentence-level approaches are ill-suited to the task and introduce a generative transformer-based encoderdecoder framework (GRIT) that is designed …

  4. 22.9 A 12nm 18.1TFLOPs/W Sparse Transformer Processor with Entropy-Based Early Exit, Mixed-Precision Predication and Fine-Grained Power Management

    2023

    Large language models have substantially advanced nuance and context understanding in natural language processing (NLP), further fueling the growth of intelligent conversational interfaces and virtual assistants. However, their hefty computational and memory demands make them …

  5. Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

    2015 · arXiv (Cornell University)

    One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue for …

  6. A Neural Attention Model for Abstractive Sentence Summarization

    2015

    Summarization based on text extraction is inherently limited, but generation-style abstractive methods have proven challenging to build. In this work, we propose a fully data-driven approach to abstractive sentence summarization. Our method utilizes a local …

  7. Character-Aware Neural Language Models

    2015 · arXiv (Cornell University)

    We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway network over characters, whose …

  8. Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution

    2015

    Sam Wiseman, Alexander M. Rush, Stuart Shieber, Jason Weston. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). …

  9. A Neural Attention Model for Sentence Summarization

    2015

    Summarization based on text extraction is inherently limited, but generation-style ab-stractive methods have proven challeng-ing to build. In this work, we propose a fully data-driven approach to abstrac-tive sentence summarization. Our method utilizes a local …

  10. Sequence-Level Knowledge Distillation

    2016 · arXiv (Cornell University)

    Neural machine translation (NMT) offers a novel alternative formulation of translation that is potentially simpler than statistical approaches. However to reach competitive performance, NMT models need to be exceedingly large. In this paper we consider …

  11. OpenNMT: Open-Source Toolkit for Neural Machine Translation

    2017

    We describe an open-source toolkit for neural machine translation (NMT). The toolkit prioritizes efficiency, modularity, and extensibility with the goal of supporting NMT research into model architectures, feature representations, and source modalities, while maintaining competitive …

  12. Challenges in Data-to-Document Generation

    2017

    Recent neural models have shown significant progress on the problem of generating short descriptive texts conditioned on a small number of database records. In this work, we suggest a slightly more difficult data-to-text generation task, …

  13. OpenNMT: Neural Machine Translation Toolkit

    2018 · arXiv (Cornell University)

    OpenNMT is an open-source toolkit for neural machine translation (NMT). The system prioritizes efficiency, modularity, and extensibility with the goal of supporting NMT research into model architectures, feature representations, and source modalities, while maintaining competitive …

  14. Bottom-Up Abstractive Summarization

    2018

    Neural network-based methods for abstractive summarization produce outputs that are more fluent than other techniques, but which can be poor at content selection. This work proposes a simple technique for addressing this issue: use a …

  15. Compound Probabilistic Context-Free Grammars for Grammar Induction

    2019

    We study a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar. In contrast to traditional formulations which learn a single stochastic grammar, our context-free …

  16. Character-Aware Neural Language Models

    2016 · Proceedings of the AAAI Conference on Artificial Intelligence

    We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway net work over characters, …

  17. Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

    2016 · International Conference on Learning Representations

    Abstract: One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue …

  18. Sequence-to-Sequence Learning as Beam-Search Optimization

    2016

    Sequence-to-Sequence (seq2seq) modeling has rapidly become an important generalpurpose NLP tool that has proven effective for many text-generation and sequence-labeling tasks. Seq2seq builds on deep neural language modeling and inherits its remarkable accuracy in estimating …

  19. Learning Global Features for Coreference Resolution

    2016

    There is compelling evidence that coreference prediction would benefit from modeling global information about entity-clusters. Yet, state-of-the-art performance can be achieved with systems treating each mention prediction independently, which we attribute to the inherent difficulty …

  20. Commonsense Knowledge Mining from Pretrained Models

    2019

    Joe Davison, Joshua Feldman, Alexander Rush. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  21. Transformers: State-of-the-Art Natural Language Processing

    2020

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, …

  22. Datasets: A Community Library for Natural Language Processing

    2021

    Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le …

  23. Multitask Prompted Training Enables Zero-Shot Task Generalization

    2021 · arXiv (Cornell University)

    Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning …

  24. PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts

    2022

    Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, …

  25. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    2022 · arXiv (Cornell University)

    Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …