ملف الباحث

Hana Hoshino

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. OPIRL: Sample Efficient Off-Policy Inverse Reinforcement Learning via Distribution Matching

    2021 · arXiv (Cornell University)

    Inverse Reinforcement Learning (IRL) is attractive in scenarios where reward engineering can be tedious. However, prior IRL algorithms use on-policy transitions, which require intensive sampling from the current policy for stable and optimal performance. This …