ملف الباحث
Weipeng Liu
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Use advantage-better action imitation for policy constraint
2023
Offline Reinforcement Learning (RL) aims to learn an optimal policy from a fixed dataset previously collected. Unlike in the online training process, the errors in value estimation from out-of-distribution actions (OOD actions) could not be …