Canwen Xu
7 papers in the PaperMetrix corpus
Papers by this author
-
InforMask: Unsupervised Informative Masking for Language Model Pretraining
2022 · arXiv (Cornell University)
Masked language modeling is widely used for pretraining large language models for natural language understanding (NLU). However, random masking is suboptimal, allocating an equal masking rate for all tokens. In this paper, we propose InforMask, …
-
Automatic Pair Construction for Contrastive Post-training
2024
Canwen Xu, Corby Rosset, Ethan Chau, Luciano Corro, Shweti Mahajan, Julian McAuley, Jennifer Neville, Ahmed Awadallah, Nikhil Rao. Findings of the Association for Computational Linguistics: NAACL 2024. 2024.
-
Transformers: State-of-the-Art Natural Language Processing
2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, …
-
Datasets: A Community Library for Natural Language Processing
2021
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le …
-
Multitask Prompted Training Enables Zero-Shot Task Generalization
2021 · arXiv (Cornell University)
Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning …
-
PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
2022
Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, …
-
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
2022 · arXiv (Cornell University)
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …