Alexander M. Rush
25 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
On Adversarial Removal of Hypothesis-only Bias in Natural Language Inference
2019
Popular Natural Language Inference (NLI) datasets have been shown to be tainted by hypothesis-only biases. Adversarial learning may help models ignore sensitive biases and spurious correlations in data. We evaluate whether adversarial learning can be …
-
Sequence-Level Mixed Sample Data Augmentation
2020
Despite their empirical success, neural networks still have difficulty capturing compositional aspects of natural language. This work proposes a simple data augmentation approach to encourage compositional behavior in neural models for sequence-to-sequence problems. Our approach, …
-
GRIT: Generative Role-filler Transformers for Document-level Event Entity Extraction
2021
We revisit the classic problem of documentlevel role-filler entity extraction (REE) for template filling. We argue that sentence-level approaches are ill-suited to the task and introduce a generative transformer-based encoderdecoder framework (GRIT) that is designed …
-
22.9 A 12nm 18.1TFLOPs/W Sparse Transformer Processor with Entropy-Based Early Exit, Mixed-Precision Predication and Fine-Grained Power Management
2023
Large language models have substantially advanced nuance and context understanding in natural language processing (NLP), further fueling the growth of intelligent conversational interfaces and virtual assistants. However, their hefty computational and memory demands make them …
-
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
2015 · arXiv (Cornell University)
One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue for …
-
A Neural Attention Model for Abstractive Sentence Summarization
2015
Summarization based on text extraction is inherently limited, but generation-style abstractive methods have proven challenging to build. In this work, we propose a fully data-driven approach to abstractive sentence summarization. Our method utilizes a local …
-
Character-Aware Neural Language Models
2015 · arXiv (Cornell University)
We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway network over characters, whose …
-
Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution
2015
Sam Wiseman, Alexander M. Rush, Stuart Shieber, Jason Weston. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). …
-
A Neural Attention Model for Sentence Summarization
2015
Summarization based on text extraction is inherently limited, but generation-style ab-stractive methods have proven challeng-ing to build. In this work, we propose a fully data-driven approach to abstrac-tive sentence summarization. Our method utilizes a local …
-
Sequence-Level Knowledge Distillation
2016 · arXiv (Cornell University)
Neural machine translation (NMT) offers a novel alternative formulation of translation that is potentially simpler than statistical approaches. However to reach competitive performance, NMT models need to be exceedingly large. In this paper we consider …
-
OpenNMT: Open-Source Toolkit for Neural Machine Translation
2017
We describe an open-source toolkit for neural machine translation (NMT). The toolkit prioritizes efficiency, modularity, and extensibility with the goal of supporting NMT research into model architectures, feature representations, and source modalities, while maintaining competitive …
-
Challenges in Data-to-Document Generation
2017
Recent neural models have shown significant progress on the problem of generating short descriptive texts conditioned on a small number of database records. In this work, we suggest a slightly more difficult data-to-text generation task, …
-
OpenNMT: Neural Machine Translation Toolkit
2018 · arXiv (Cornell University)
OpenNMT is an open-source toolkit for neural machine translation (NMT). The system prioritizes efficiency, modularity, and extensibility with the goal of supporting NMT research into model architectures, feature representations, and source modalities, while maintaining competitive …
-
Bottom-Up Abstractive Summarization
2018
Neural network-based methods for abstractive summarization produce outputs that are more fluent than other techniques, but which can be poor at content selection. This work proposes a simple technique for addressing this issue: use a …
-
Compound Probabilistic Context-Free Grammars for Grammar Induction
2019
We study a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar. In contrast to traditional formulations which learn a single stochastic grammar, our context-free …
-
Character-Aware Neural Language Models
2016 · Proceedings of the AAAI Conference on Artificial Intelligence
We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway net work over characters, …
-
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
2016 · International Conference on Learning Representations
Abstract: One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue …
-
Sequence-to-Sequence Learning as Beam-Search Optimization
2016
Sequence-to-Sequence (seq2seq) modeling has rapidly become an important generalpurpose NLP tool that has proven effective for many text-generation and sequence-labeling tasks. Seq2seq builds on deep neural language modeling and inherits its remarkable accuracy in estimating …
-
Learning Global Features for Coreference Resolution
2016
There is compelling evidence that coreference prediction would benefit from modeling global information about entity-clusters. Yet, state-of-the-art performance can be achieved with systems treating each mention prediction independently, which we attribute to the inherent difficulty …
-
Commonsense Knowledge Mining from Pretrained Models
2019
Joe Davison, Joshua Feldman, Alexander Rush. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
Transformers: State-of-the-Art Natural Language Processing
2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, …
-
Datasets: A Community Library for Natural Language Processing
2021
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le …
-
Multitask Prompted Training Enables Zero-Shot Task Generalization
2021 · arXiv (Cornell University)
Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning …
-
PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
2022
Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, …
-
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
2022 · arXiv (Cornell University)
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …