conference-paper
Open access
HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization
Research footprint
At a glance
- Citations
- 364
- References
- 49
- Comments
- 0
Paper overview
Abstract
Neural extractive summarization models usually employ a hierarchical encoder for document encoding and they are trained using sentence-level labels, which are created heuristically using rule-based methods. Training the hierarchical encoder with these inaccurate labels is challenging. Inspired by the recent work on pre-training transformer sentence encoders We apply the pre-trained HIBERT to our summarization model and it outperforms its randomly initialized counterpart by 1.25 ROUGE on the CNN/Dailymail dataset and by 2.0 ROUGE on a version of New York Times dataset. We also achieve the state-of-the-art performance on these two datasets.
Record transparency
Publication details
- DOI
- 10.18653/v1/p19-1499
- OpenAlex
- W2962785754
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.