Align then Summarize: Automatic Alignment Methods for Summarization\n Corpus Creation
At a glance
- Citations
- 1
- References
- 0
- Comments
- 0
Öz
Summarizing texts is not a straightforward task. Before even considering text\nsummarization, one should determine what kind of summary is expected. How much\nshould the information be compressed? Is it relevant to reformulate or should\nthe summary stick to the original phrasing? State-of-the-art on automatic text\nsummarization mostly revolves around news articles. We suggest that considering\na wider variety of tasks would lead to an improvement in the field, in terms of\ngeneralization and robustness. We explore meeting summarization: generating\nreports from automatic transcriptions. Our work consists in segmenting and\naligning transcriptions with respect to reports, to get a suitable dataset for\nneural summarization. Using a bootstrapping approach, we provide pre-alignments\nthat are corrected by human annotators, making a validation set against which\nwe evaluate automatic models. This consistently reduces annotators' efforts by\nproviding iteratively better pre-alignment and maximizes the corpus size by\nusing annotations from our automatic alignment models. Evaluation is conducted\non \\publicmeetings, a novel corpus of aligned public meetings. We report\nautomatic alignment and summarization performances on this corpus and show that\nautomatic alignment is relevant for data annotation since it leads to large\nimprovement of almost +4 on all ROUGE scores on the summarization task.\n
Publication details
- DOI
- 10.48550/arxiv.2007.07841
- OpenAlex
- W4287719367
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Oturum Açın to join the discussion.