ملف الباحث

Mathieu Petitbois

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment

    2026 · HAL (Le Centre pour la Communication Scientifique Directe)

    We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions. In this setting, aligning style with high task performance is particularly challenging due to distribution shift and inherent conflicts …