ملف الباحث
Zixian Guo
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
SelfBC: Self Behavior Cloning for Offline Reinforcement Learning
2024 · arXiv (Cornell University)
Policy constraint methods in offline reinforcement learning employ additional regularization techniques to constrain the discrepancy between the learned policy and the offline dataset. However, these methods tend to result in overly conservative policies that resemble …