ملف الباحث
Steve Roberts
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Stabilizing Off-Policy Deep Reinforcement Learning from Pixels
2022 · arXiv (Cornell University)
Off-policy reinforcement learning (RL) from pixel observations is notoriously unstable. As a result, many successful algorithms must combine different domain-specific practices and auxiliary losses to learn meaningful behaviors in complex environments. In this work, we …