preprint Open access

Summarizing User-generated Textual Content: Motivation and Methods for\n Fairness in Algorithmic Summaries

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

As the amount of user-generated textual content grows rapidly, text\nsummarization algorithms are increasingly being used to provide users a quick\noverview of the information content. Traditionally, summarization algorithms\nhave been evaluated only based on how well they match human-written summaries\n(e.g. as measured by ROUGE scores). In this work, we propose to evaluate\nsummarization algorithms from a completely new perspective that is important\nwhen the user-generated data to be summarized comes from different socially\nsalient user groups, e.g. men or women, Caucasians or African-Americans, or\ndifferent political groups (Republicans or Democrats). In such cases, we check\nwhether the generated summaries fairly represent these different social groups.\nSpecifically, considering that an extractive summarization algorithm selects a\nsubset of the textual units (e.g. microblogs) in the original data for\ninclusion in the summary, we investigate whether this selection is fair or not.\nOur experiments over real-world microblog datasets show that existing\nsummarization algorithms often represent the socially salient user-groups very\ndifferently compared to their distributions in the original data. More\nimportantly, some groups are frequently under-represented in the generated\nsummaries, and hence get far less exposure than what they would have obtained\nin the original data. To reduce such adverse impacts, we propose novel\nfairness-preserving summarization algorithms which produce high-quality\nsummaries while ensuring fairness among various groups. To our knowledge, this\nis the first attempt to produce fair text summarization, and is likely to open\nup an interesting research direction.\n

Record transparency

Publication details

DOI
10.48550/arxiv.1810.09147
OpenAlex
W4289373227
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.