Mihir Kale
5 papers in the PaperMetrix corpus
Papers by this author
-
nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?
2021 · arXiv (Cornell University)
Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of …
-
mT5: A massively multilingual pre-trained text-to-text transformer
2020 · arXiv (Cornell University)
The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …
-
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
2022 · Transactions of the Association for Computational Linguistics
Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …
-
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
2021
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …
-
The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
2021
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, …