Marlos C. Machado
4 papers in the PaperMetrix corpus
Papers by this author
-
Learning Purposeful Behaviour in the Absence of Rewards
2016 · arXiv (Cornell University)
Artificial intelligence is commonly defined as the ability to achieve goals in the world. In the reinforcement learning framework, goals are encoded as reward functions that guide agent behaviour, and the sum of observed rewards …
-
A Laplacian Framework for Option Discovery in Reinforcement Learning
2017 · arXiv (Cornell University)
Representation learning and option discovery are two of the biggest challenges in reinforcement learning (RL). Proto-value functions (PVFs) are a well-known approach for representation learning in MDPs. In this paper we address the option discovery …
-
A functional mirror ascent view of policy gradient methods with function approximation.
2021 · arXiv (Cornell University)
We use functional mirror ascent to propose a general framework (referred to as FMA-PG) for designing policy gradient methods. The functional perspective distinguishes between a policy's functional representation (what are its sufficient statistics) and its …
-
Deep Reinforcement Learning with Gradient Eligibility Traces
2025 · arXiv (Cornell University)
Achieving fast and stable off-policy learning in deep reinforcement learning (RL) is challenging. Most existing methods rely on semi-gradient temporal-difference (TD) methods for their simplicity and efficiency, but are consequently susceptible to divergence. While more …