Researcher profile

Mihir Kale

5 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?

    2021 · arXiv (Cornell University)

    Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of …

  2. mT5: A massively multilingual pre-trained text-to-text transformer

    2020 · arXiv (Cornell University)

    The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …

  3. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models

    2022 · Transactions of the Association for Computational Linguistics

    Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …

  4. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

    2021

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …

  5. The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

    2021

    Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, …