Researcher profile
Harshat Kumar
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Actor-only Deterministic Policy Gradient via Zeroth-order Gradient Oracles in Action Space
2021
Deterministic policies demonstrate substantial empirical success over their stochastic counterparts as they remove a level of randomness in Policy Gradient (PG) methods when applied to stochastic search problems involving Markov decision processes. However, current implementations …
-
On the Sample Complexity of Actor-Critic Method for Reinforcement Learning with Function Approximation
2019 · arXiv (Cornell University)
Reinforcement learning, mathematically described by Markov Decision Problems, may be approached either through dynamic programming or policy search. Actor-critic algorithms combine the merits of both approaches by alternating between steps to estimate the value function …