conference-paper Open access

MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization

Research footprint

At a glance

Citations
90
References
21
Comments
0
Paper overview

Abstract

This paper introduces MEDIASUM 1 , a largescale media interview dataset consisting of 463.6K transcripts with abstractive summaries. To create this dataset, we collect interview transcripts from NPR and CNN and employ the overview and topic descriptions as summaries. Compared with existing public corpora for dialogue summarization, our dataset is an order of magnitude larger and contains complex multi-party conversations from multiple domains. We conduct statistical analysis to demonstrate the unique positional bias exhibited in the transcripts of televised and radioed interviews. We also show that MEDIASUM can be used in transfer learning to improve a model's performance on other dialogue summarization tasks.

Record transparency

Publication details

DOI
10.18653/v1/2021.naacl-main.474
OpenAlex
W3169942382
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.