conference-paper Open access

POLICEd RL: Learning Closed-Loop Robot Control Policies with Provable Satisfaction of Hard Constraints

Research footprint

At a glance

Citations
4
References
0
Comments
0
Paper overview

Öz

In this paper, we seek to learn a robot policy guaranteed to satisfy state constraints.To encourage constraint satisfaction, existing RL algorithms typically rely on Constrained Markov Decision Processes and discourage constraint violations through reward shaping.However, such soft constraints cannot offer safety guarantees.To address this gap, we propose POLICEd RL, a novel RL algorithm explicitly designed to enforce affine hard constraints in closed-loop with a black-box environment.Our key insight is to make the learned policy be affine around the unsafe set and to use this affine region as a repulsive buffer to prevent trajectories from violating the constraint.We prove that such policies exist and guarantee constraint satisfaction.Our proposed framework is applicable to both systems with continuous and discrete state and action spaces and is agnostic to the choice of the RL training algorithm.Our results demonstrate the capacity of POLICEd RL to enforce hard constraints in robotic tasks while significantly outperforming existing methods.

Record transparency

Publication details

DOI
10.15607/rss.2024.xx.104
OpenAlex
W4402353995
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.