article Open access

A method to utilize prior knowledge for extractive summarization based on pre-trained language models

  • Vietnam Journal of Science and Technology/Science and Technology
  • Vietnam Academy of Science and Technology
Research footprint

At a glance

Citations
0
References
29
Comments
0
Paper overview

Abstract

This paper presents a novel model for extractive summarization that integrates context representation from a pre-trained language model (PLM), such as BERT, with prior knowledge derived from unsupervised learning methods. Sentence importance assessment is crucial in extractive summarization, with prior knowledge providing indicators of sentence importance within a document. Our model introduces a method for estimating sentence importance based on prior knowledge, complementing the contextual representation offered by PLMs like BERT. Unlike previous approaches that primarily relied on PLMs alone, our model leverages both contextual representation and prior knowledge extracted from each input document. By conditioning the model on prior knowledge, it emphasizes key sentences in generating the final summary. We evaluate our model on three benchmark datasets across two languages, demonstrating improved performance compared to strong baseline methods in extractive summarization. Additionally, our ablation study reveals that injecting knowledge into certain first attention layers yields greater benefits than others. The model code is publicly available for further exploration.

Record transparency

Publication details

DOI
10.15625/2525-2518/20241
OpenAlex
W4405088531
Document type
article
Language
EN
Source
Vietnam Journal of Science and Technology/Science and Technology
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.