Researcher profile

Rami Al‐Rfou

8 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Watch Your Step: Learning Node Embeddings via Graph Attention

    2017 · arXiv (Cornell University)

    Graph embedding methods represent nodes in a continuous vector space, preserving information from the graph (e.g. by sampling random walks). There are many hyper-parameters to these methods (such as random walk length) which have to …

  2. nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?

    2021 · arXiv (Cornell University)

    Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of …

  3. Statistically Significant Detection of Linguistic Change

    2015

    We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the rapid exchange of …

  4. Character-Level Language Modeling with Deeper Self-Attention

    2019

    LSTMs and other RNN variants have shown strong performance on character-level language modeling. These models are typically trained using truncated backpropagation through time, and it is common to assume that their success stems from their …

  5. mT5: A massively multilingual pre-trained text-to-text transformer

    2020 · arXiv (Cornell University)

    The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …

  6. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models

    2022 · Transactions of the Association for Computational Linguistics

    Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …

  7. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

    2021

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …

  8. SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer

    2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    There has been growing interest in parameterefficient methods to apply pre-trained language models to downstream tasks. Building on the PROMPTTUNING approach of Lester et al. ( SPOT first learns a prompt on one or more …