conference-paper

Actor-only Deterministic Policy Gradient via Zeroth-order Gradient Oracles in Action Space

Research footprint

At a glance

Citations
1
References
52
Comments
0
Paper overview

Abstract

Deterministic policies demonstrate substantial empirical success over their stochastic counterparts as they remove a level of randomness in Policy Gradient (PG) methods when applied to stochastic search problems involving Markov decision processes. However, current implementations require the use of state-action value ($Q$-function) approximators, also known as critics, to obtain estimates of the associated policy-reward gradient. In this work, we propose the use of two-point stochastic evaluations to obtain gradient estimates of a smoothed$Q$-function surrogate, constructed by evaluating pairs of the$Q$-function at low-dimensional, randomized initial action perturbations. This procedure lifts the dependence on a critic and restores true model-free policy learning, and with provable algorithmic stability. In fact, our finite complexity bounds improve upon existing results by up to 2 orders of magnitude in terms of iteration complexity, and by up to 3/2 orders of magnitude in terms of sample complexity. Simulation results on an agent navigation problem showcase the effectiveness of our proposed algorithm in a practical setting, as well.

Record transparency

Publication details

DOI
10.1109/isit45174.2021.9518023
OpenAlex
W3198568807
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.