Martha White
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
2020 · arXiv (Cornell University)
Dyna-style reinforcement learning (RL) agents improve sample efficiency over model-free RL agents by updating the value function with simulated experience generated by an environment model. However, it is often difficult to learn accurate models of …
-
Adapting Behavior via Intrinsic Reward: A Survey and Empirical Study
2020 · Journal of Artificial Intelligence Research
Learning about many things can provide numerous benefits to a reinforcement learning system. For example, learning many auxiliary value functions, in addition to optimizing the environmental reward, appears to improve both exploration and representation learning. …
-
Deep Reinforcement Learning with Gradient Eligibility Traces
2025 · arXiv (Cornell University)
Achieving fast and stable off-policy learning in deep reinforcement learning (RL) is challenging. Most existing methods rely on semi-gradient temporal-difference (TD) methods for their simplicity and efficiency, but are consequently susceptible to divergence. While more …