Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
At a glance
- Citations
- 3
- References
- 65
- Comments
- 0
Öz
Artificial Intelligence, particularly through recent advancements in deep learning (DL), has achieved exceptional performances in many tasks in fields such as natural language processing and computer vision. For certain high-stake domains, in addition to desirable performance metrics, a high level of interpretability is often required in order for AI to be reliably utilized. Unfortunately, the black box nature of DL models prevents researchers from providing explicative descriptions for a DL model’s reasoning process and decisions. In this work, we propose a novel framework utilizing Adversarial Inverse Reinforcement Learning that can provide global explanations for decisions made by a Reinforcement Learning model and capture intuitive tendencies that the model follows by summarizing the model’s decision-making process.
Publication details
- DOI
- 10.1109/icpr56361.2022.9956245
- OpenAlex
- W4312871112
- Document type
- conference-paper
- Language
- EN
- Source
- 2022 26th International Conference on Pattern Recognition (ICPR)
- Last metadata update
Comments
Oturum Açın to join the discussion.