Actor-only Deterministic Policy Gradient via Zeroth-order Gradient Oracles in Action Space
At a glance
- Citations
- 1
- References
- 52
- Comments
- 0
Öz
Deterministic policies demonstrate substantial empirical success over their stochastic counterparts as they remove a level of randomness in Policy Gradient (PG) methods when applied to stochastic search problems involving Markov decision processes. However, current implementations require the use of state-action value ($Q$-function) approximators, also known as critics, to obtain estimates of the associated policy-reward gradient. In this work, we propose the use of two-point stochastic evaluations to obtain gradient estimates of a smoothed$Q$-function surrogate, constructed by evaluating pairs of the$Q$-function at low-dimensional, randomized initial action perturbations. This procedure lifts the dependence on a critic and restores true model-free policy learning, and with provable algorithmic stability. In fact, our finite complexity bounds improve upon existing results by up to 2 orders of magnitude in terms of iteration complexity, and by up to 3/2 orders of magnitude in terms of sample complexity. Simulation results on an agent navigation problem showcase the effectiveness of our proposed algorithm in a practical setting, as well.
Publication details
- DOI
- 10.1109/isit45174.2021.9518023
- OpenAlex
- W3198568807
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.