Colin Raffel
15 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Deduplicating Training Data Mitigates Privacy Risks in Language Models
2022 · arXiv (Cornell University)
Past work has shown that large language models are susceptible to privacy attacks, where adversaries generate sequences from a trained model and detect which sequences are memorized from the training set. In this work, we …
-
Emergent Abilities of Large Language Models
2022 · arXiv (Cornell University)
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities …
-
Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models
2023 · arXiv (Cornell University)
Currently, most machine learning models are trained by centralized teams and are rarely updated. In contrast, open-source software development involves the iterative development of a shared artifact through distributed collaboration using a version control system. …
-
Merging by Matching Models in Task Parameter Subspaces
2023 · arXiv (Cornell University)
Model merging aims to cheaply combine individual task-specific models into a single multitask model. In this work, we view past merging methods as leveraging different notions of a ''task parameter subspace'' in which models are …
-
Online and Linear-Time Attention by Enforcing Monotonic Alignments
2017 · arXiv (Cornell University)
Recurrent neural network models with an attention mechanism have proven to be extremely effective on a wide variety of sequence-to-sequence problems. However, the fact that soft attention mechanisms perform a pass over the entire input …
-
Monotonic Chunkwise Attention
2017 · arXiv (Cornell University)
Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence transduction. To address …
-
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
2019 · arXiv (Cornell University)
Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …
-
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
2019 · arXiv (Cornell University)
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning …
-
mT5: A massively multilingual pre-trained text-to-text transformer
2020 · arXiv (Cornell University)
The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …
-
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
2022 · Transactions of the Association for Computational Linguistics
Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …
-
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
2021
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …
-
Multitask Prompted Training Enables Zero-Shot Task Generalization
2021 · arXiv (Cornell University)
Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning …
-
PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
2022
Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, …
-
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
2022 · arXiv (Cornell University)
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …
-
Crosslingual Generalization through Multitask Finetuning
2023
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid …