conference-paper Open access

LongT5: Efficient Text-To-Text Transformer for Long Sequences

  • Findings of the Association for Computational Linguistics: NAACL 2022
Research footprint

At a glance

Citations
216
References
39
Comments
0
Paper overview

Öz

Recent work has shown that either (1) increasing the input length or (2) increasing model size can improve the performance of Transformer-based neural models. In this paper, we present LongT5, a new model that explores the effects of scaling both the input length and model size at the same time. Specifically, we integrate attention ideas from long-input transformers (ETC), and adopt pretraining strategies from summarization pretraining (PEGASUS) into the scalable T5 architecture. The result is a new attention mechanism we call Transient Global (TGlobal), which mimics ETC's local/global attention mechanism, but without requiring additional side-inputs. We are able to achieve state-ofthe-art results on several summarization and question answering tasks, as well as outperform the original T5 models on these tasks. We have open sourced our architecture and training code, as well as our pre-trained model checkpoints.

Record transparency

Publication details

DOI
10.18653/v1/2022.findings-naacl.55
OpenAlex
W4225727438
Document type
conference-paper
Language
EN
Source
Findings of the Association for Computational Linguistics: NAACL 2022
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.