preprint Open access

The Explanation Game: Towards Prediction Explainability through Sparse\n Communication

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
1
References
0
Comments
0
Paper overview

Abstract

Explainability is a topic of growing importance in NLP. In this work, we\nprovide a unified perspective of explainability as a communication problem\nbetween an explainer and a layperson about a classifier's decision. We use this\nframework to compare several prior approaches for extracting explanations,\nincluding gradient methods, representation erasure, and attention mechanisms,\nin terms of their communication success. In addition, we reinterpret these\nmethods at the light of classical feature selection, and we use this as\ninspiration to propose new embedded methods for explainability, through the use\nof selective, sparse attention. Experiments in text classification, natural\nlanguage entailment, and machine translation, using different configurations of\nexplainers and laypeople (including both machines and humans), reveal an\nadvantage of attention-based explainers over gradient and erasure methods.\nFurthermore, human evaluation experiments show promising results with post-hoc\nexplainers trained to optimize communication success and faithfulness.\n

Record transparency

Publication details

DOI
10.48550/arxiv.2004.13876
OpenAlex
W4287811189
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.