Matthijs T. J. Spaan
4 papers in the PaperMetrix corpus
Papers by this author
-
Bounding the Probability of Resource Constraint Violations in Multi-Agent MDPs
2017 · Proceedings of the AAAI Conference on Artificial Intelligence
Multi-agent planning problems with constraints on global resource consumption occur in several domains. Existing algorithms for solving Multi-agent Markov Decision Processes can compute policies that meet a resource constraint in expectation, but these policies provide …
-
Structure Learning for Safe Policy Improvement
2019
We investigate how Safe Policy Improvement (SPI) algorithms can exploit the structure of factored Markov decision processes when such structure is unknown a priori. To facilitate the application of reinforcement learning in the real world, …
-
CEM: Constrained Entropy Maximization for Task-Agnostic Safe Exploration
2023 · Proceedings of the AAAI Conference on Artificial Intelligence
In the absence of assigned tasks, a learning agent typically seeks to explore its environment efficiently. However, the pursuit of exploration will bring more safety risks. An under-explored aspect of reinforcement learning is how to …
-
Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs
2024 · arXiv (Cornell University)
In the zero-shot policy transfer (ZSPT) setting for contextual Markov decision processes (CMDP), agents train on a fixed, finite set of contexts and must generalize to new ones. Recent work has demonstrated that training on …