conference-paper

Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning

  • 2022 26th International Conference on Pattern Recognition (ICPR)
Research footprint

At a glance

Citations
3
References
65
Comments
0
Paper overview

Abstract

Artificial Intelligence, particularly through recent advancements in deep learning (DL), has achieved exceptional performances in many tasks in fields such as natural language processing and computer vision. For certain high-stake domains, in addition to desirable performance metrics, a high level of interpretability is often required in order for AI to be reliably utilized. Unfortunately, the black box nature of DL models prevents researchers from providing explicative descriptions for a DL model’s reasoning process and decisions. In this work, we propose a novel framework utilizing Adversarial Inverse Reinforcement Learning that can provide global explanations for decisions made by a Reinforcement Learning model and capture intuitive tendencies that the model follows by summarizing the model’s decision-making process.

Record transparency

Publication details

DOI
10.1109/icpr56361.2022.9956245
OpenAlex
W4312871112
Document type
conference-paper
Language
EN
Source
2022 26th International Conference on Pattern Recognition (ICPR)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.