Researcher profile

Linting Xue

4 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?

    2021 · arXiv (Cornell University)

    Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of …

  2. mT5: A massively multilingual pre-trained text-to-text transformer

    2020 · arXiv (Cornell University)

    The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …

  3. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models

    2022 · Transactions of the Association for Computational Linguistics

    Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …

  4. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

    2021

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …