conference-paper Open access

HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization

Research footprint

At a glance

Citations
364
References
49
Comments
0
Paper overview

Abstract

Neural extractive summarization models usually employ a hierarchical encoder for document encoding and they are trained using sentence-level labels, which are created heuristically using rule-based methods. Training the hierarchical encoder with these inaccurate labels is challenging. Inspired by the recent work on pre-training transformer sentence encoders We apply the pre-trained HIBERT to our summarization model and it outperforms its randomly initialized counterpart by 1.25 ROUGE on the CNN/Dailymail dataset and by 2.0 ROUGE on a version of New York Times dataset. We also achieve the state-of-the-art performance on these two datasets.

Record transparency

Publication details

DOI
10.18653/v1/p19-1499
OpenAlex
W2962785754
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.