The Explanation Game: Towards Prediction Explainability through Sparse\n Communication
At a glance
- Citations
- 1
- References
- 0
- Comments
- 0
Abstract
Explainability is a topic of growing importance in NLP. In this work, we\nprovide a unified perspective of explainability as a communication problem\nbetween an explainer and a layperson about a classifier's decision. We use this\nframework to compare several prior approaches for extracting explanations,\nincluding gradient methods, representation erasure, and attention mechanisms,\nin terms of their communication success. In addition, we reinterpret these\nmethods at the light of classical feature selection, and we use this as\ninspiration to propose new embedded methods for explainability, through the use\nof selective, sparse attention. Experiments in text classification, natural\nlanguage entailment, and machine translation, using different configurations of\nexplainers and laypeople (including both machines and humans), reveal an\nadvantage of attention-based explainers over gradient and erasure methods.\nFurthermore, human evaluation experiments show promising results with post-hoc\nexplainers trained to optimize communication success and faithfulness.\n
Publication details
- DOI
- 10.48550/arxiv.2004.13876
- OpenAlex
- W4287811189
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.