ملف الباحث
Hepeng Li
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
An Improved Trust-Region Method for Off-Policy Deep Reinforcement Learning
2023
Reinforcement learning (RL) is a powerful tool for training agents to interact with complex environments. In particular, trust-region methods are widely used for policy optimization in model-free RL. However, these methods suffer from high sample …