conference-paper

Improving Summarization Quality with Topic Modeling

Research footprint

At a glance

Citations
3
References
34
Comments
0
Paper overview

Abstract

The problem of extractive text summarization for a collection of documents is defined as the problem of selecting a small subset of sentences so that the contents and meaning of the original document set are preserved in the best possible way. In this paper we describe different applications of topic modeling as it relates to summarization. We consider two summarization models - supervised and unsupervised - enhanced by integrating the topic knowledge into their standard lexicon-based models. Both summarizers strive to cover as much information of the input documents as possible when generating the summaries. The supervised summarizer operates in a standard rank-and-select-sentences manner of extractive summarization, where the best linear combination of multiple sentence features is learned by a genetic algorithm (GA). The unsupervised summarizer models the summarization task as an optimization problem. As is the case with most existing summarization approaches, both original models measure information coverage by lexical units and were enriched by topic knowledge, providing a new measure for the informaition coverage. The experimental results show that utilizing topic knowledge improves the summarization quality.

Record transparency

Publication details

DOI
10.1145/2809936.2809944
OpenAlex
W2070120463
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.