ملف الباحث

Weipeng Liu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Use advantage-better action imitation for policy constraint

    2023

    Offline Reinforcement Learning (RL) aims to learn an optimal policy from a fixed dataset previously collected. Unlike in the online training process, the errors in value estimation from out-of-distribution actions (OOD actions) could not be …