Rami Al‐Rfou
8 papers in the PaperMetrix corpus
Papers by this author
-
Watch Your Step: Learning Node Embeddings via Graph Attention
2017 · arXiv (Cornell University)
Graph embedding methods represent nodes in a continuous vector space, preserving information from the graph (e.g. by sampling random walks). There are many hyper-parameters to these methods (such as random walk length) which have to …
-
nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?
2021 · arXiv (Cornell University)
Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of …
-
Statistically Significant Detection of Linguistic Change
2015
We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the rapid exchange of …
-
Character-Level Language Modeling with Deeper Self-Attention
2019
LSTMs and other RNN variants have shown strong performance on character-level language modeling. These models are typically trained using truncated backpropagation through time, and it is common to assume that their success stems from their …
-
mT5: A massively multilingual pre-trained text-to-text transformer
2020 · arXiv (Cornell University)
The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …
-
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
2022 · Transactions of the Association for Computational Linguistics
Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …
-
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
2021
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …
-
SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer
2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
There has been growing interest in parameterefficient methods to apply pre-trained language models to downstream tasks. Building on the PROMPTTUNING approach of Lester et al. ( SPOT first learns a prompt on one or more …