Researcher profile
Maolin Hou
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Use advantage-better action imitation for policy constraint
2023
Offline Reinforcement Learning (RL) aims to learn an optimal policy from a fixed dataset previously collected. Unlike in the online training process, the errors in value estimation from out-of-distribution actions (OOD actions) could not be …