Researcher profile

Matthijs T. J. Spaan

4 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Bounding the Probability of Resource Constraint Violations in Multi-Agent MDPs

    2017 · Proceedings of the AAAI Conference on Artificial Intelligence

    Multi-agent planning problems with constraints on global resource consumption occur in several domains. Existing algorithms for solving Multi-agent Markov Decision Processes can compute policies that meet a resource constraint in expectation, but these policies provide …

  2. Structure Learning for Safe Policy Improvement

    2019

    We investigate how Safe Policy Improvement (SPI) algorithms can exploit the structure of factored Markov decision processes when such structure is unknown a priori. To facilitate the application of reinforcement learning in the real world, …

  3. CEM: Constrained Entropy Maximization for Task-Agnostic Safe Exploration

    2023 · Proceedings of the AAAI Conference on Artificial Intelligence

    In the absence of assigned tasks, a learning agent typically seeks to explore its environment efficiently. However, the pursuit of exploration will bring more safety risks. An under-explored aspect of reinforcement learning is how to …

  4. Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs

    2024 · arXiv (Cornell University)

    In the zero-shot policy transfer (ZSPT) setting for contextual Markov decision processes (CMDP), agents train on a fixed, finite set of contexts and must generalize to new ones. Recent work has demonstrated that training on …