ملف الباحث

Colin Raffel

15 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Deduplicating Training Data Mitigates Privacy Risks in Language Models

    2022 · arXiv (Cornell University)

    Past work has shown that large language models are susceptible to privacy attacks, where adversaries generate sequences from a trained model and detect which sequences are memorized from the training set. In this work, we …

  2. Emergent Abilities of Large Language Models

    2022 · arXiv (Cornell University)

    Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities …

  3. Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models

    2023 · arXiv (Cornell University)

    Currently, most machine learning models are trained by centralized teams and are rarely updated. In contrast, open-source software development involves the iterative development of a shared artifact through distributed collaboration using a version control system. …

  4. Merging by Matching Models in Task Parameter Subspaces

    2023 · arXiv (Cornell University)

    Model merging aims to cheaply combine individual task-specific models into a single multitask model. In this work, we view past merging methods as leveraging different notions of a ''task parameter subspace'' in which models are …

  5. Online and Linear-Time Attention by Enforcing Monotonic Alignments

    2017 · arXiv (Cornell University)

    Recurrent neural network models with an attention mechanism have proven to be extremely effective on a wide variety of sequence-to-sequence problems. However, the fact that soft attention mechanisms perform a pass over the entire input …

  6. Monotonic Chunkwise Attention

    2017 · arXiv (Cornell University)

    Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence transduction. To address …

  7. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  8. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

    2019 · arXiv (Cornell University)

    Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning …

  9. mT5: A massively multilingual pre-trained text-to-text transformer

    2020 · arXiv (Cornell University)

    The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …

  10. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models

    2022 · Transactions of the Association for Computational Linguistics

    Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …

  11. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

    2021

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …

  12. Multitask Prompted Training Enables Zero-Shot Task Generalization

    2021 · arXiv (Cornell University)

    Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning …

  13. PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts

    2022

    Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, …

  14. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    2022 · arXiv (Cornell University)

    Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …

  15. Crosslingual Generalization through Multitask Finetuning

    2023

    Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid …