Nitish Shirish Keskar
8 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Coarse-grain Fine-grain Coattention Network for Multi-evidence Question Answering
2019 · International Conference on Learning Representations
End-to-end neural models have made significant progress in question answering, however recent studies show that these models implicitly assume that the answer and evidence appear close together in a single document. In this work, we …
-
GeDi: Generative Discriminator Guided Sequence Generation
2020 · arXiv (Cornell University)
While large-scale language models (LMs) are able to imitate the distribution of natural language well enough to generate realistic text, it is difficult to control which regions of the distribution they generate. This is especially …
-
Modeling Multi-hop Question Answering as Single Sequence Prediction
2022 · arXiv (Cornell University)
Fusion-in-decoder (Fid) (Izacard and Grave, 2020) is a generative question answering (QA) model that leverages passage retrieval with a pre-trained transformer and pushed the state of the art on single-hop QA. However, the complexity of …
-
An Analysis of Neural Language Modeling at Multiple Scales
2018 · arXiv (Cornell University)
Many of the leading approaches in language modeling introduce novel, complex and specialized architectures. We take existing state-of-the-art word level language models based on LSTMs and QRNNs and extend them to both larger vocabularies as …
-
The Natural Language Decathlon: Multitask Learning as Question Answering
2018 · arXiv (Cornell University)
Deep learning has improved performance on many natural language processing (NLP) tasks individually. However, general NLP models cannot emerge within a paradigm that focuses on the particularities of a single metric, dataset, and task. We …
-
XLDA: Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering
2019 · arXiv (Cornell University)
While natural language processing systems often focus on a single language, multilingual transfer learning has the potential to improve performance, especially for low-resource languages. We introduce XLDA, cross-lingual data augmentation, a method that replaces a …
-
CTRL: A Conditional Transformer Language Model for Controllable Generation
2019 · arXiv (Cornell University)
Large-scale language models show promising text generation capabilities, but users cannot easily control particular aspects of the generated text. We release CTRL, a 1.63 billion-parameter conditional transformer language model, trained to condition on control codes …
-
GeDi: Generative Discriminator Guided Sequence Generation
2021
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, Nazneen Fatema Rajani. Findings of the Association for Computational Linguistics: EMNLP 2021. 2021.