ملف الباحث
Imani, Ehsan
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
2020 · arXiv (Cornell University)
Dyna-style reinforcement learning (RL) agents improve sample efficiency over model-free RL agents by updating the value function with simulated experience generated by an environment model. However, it is often difficult to learn accurate models of …