Improving Summarization Quality with Topic Modeling
At a glance
- Citations
- 3
- References
- 34
- Comments
- 0
Abstract
The problem of extractive text summarization for a collection of documents is defined as the problem of selecting a small subset of sentences so that the contents and meaning of the original document set are preserved in the best possible way. In this paper we describe different applications of topic modeling as it relates to summarization. We consider two summarization models - supervised and unsupervised - enhanced by integrating the topic knowledge into their standard lexicon-based models. Both summarizers strive to cover as much information of the input documents as possible when generating the summaries. The supervised summarizer operates in a standard rank-and-select-sentences manner of extractive summarization, where the best linear combination of multiple sentence features is learned by a genetic algorithm (GA). The unsupervised summarizer models the summarization task as an optimization problem. As is the case with most existing summarization approaches, both original models measure information coverage by lexical units and were enriched by topic knowledge, providing a new measure for the informaition coverage. The experimental results show that utilizing topic knowledge improves the summarization quality.
Publication details
- DOI
- 10.1145/2809936.2809944
- OpenAlex
- W2070120463
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.