preprint Open access

Align then Summarize: Automatic Alignment Methods for Summarization\n Corpus Creation

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
1
References
0
Comments
0
Paper overview

Öz

Summarizing texts is not a straightforward task. Before even considering text\nsummarization, one should determine what kind of summary is expected. How much\nshould the information be compressed? Is it relevant to reformulate or should\nthe summary stick to the original phrasing? State-of-the-art on automatic text\nsummarization mostly revolves around news articles. We suggest that considering\na wider variety of tasks would lead to an improvement in the field, in terms of\ngeneralization and robustness. We explore meeting summarization: generating\nreports from automatic transcriptions. Our work consists in segmenting and\naligning transcriptions with respect to reports, to get a suitable dataset for\nneural summarization. Using a bootstrapping approach, we provide pre-alignments\nthat are corrected by human annotators, making a validation set against which\nwe evaluate automatic models. This consistently reduces annotators' efforts by\nproviding iteratively better pre-alignment and maximizes the corpus size by\nusing annotations from our automatic alignment models. Evaluation is conducted\non \\publicmeetings, a novel corpus of aligned public meetings. We report\nautomatic alignment and summarization performances on this corpus and show that\nautomatic alignment is relevant for data annotation since it leads to large\nimprovement of almost +4 on all ROUGE scores on the summarization task.\n

Record transparency

Publication details

DOI
10.48550/arxiv.2007.07841
OpenAlex
W4287719367
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.