preprint Open access

An Efficient Asynchronous Method for Integrating Evolutionary and\n Gradient-based Policy Search

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
3
References
0
Comments
0
Paper overview

Abstract

Deep reinforcement learning (DRL) algorithms and evolution strategies (ES)\nhave been applied to various tasks, showing excellent performances. These have\nthe opposite properties, with DRL having good sample efficiency and poor\nstability, while ES being vice versa. Recently, there have been attempts to\ncombine these algorithms, but these methods fully rely on synchronous update\nscheme, making it not ideal to maximize the benefits of the parallelism in ES.\nTo solve this challenge, asynchronous update scheme was introduced, which is\ncapable of good time-efficiency and diverse policy exploration. In this paper,\nwe introduce an Asynchronous Evolution Strategy-Reinforcement Learning (AES-RL)\nthat maximizes the parallel efficiency of ES and integrates it with policy\ngradient methods. Specifically, we propose 1) a novel framework to merge ES and\nDRL asynchronously and 2) various asynchronous update methods that can take all\nadvantages of asynchronism, ES, and DRL, which are exploration and time\nefficiency, stability, and sample efficiency, respectively. The proposed\nframework and update methods are evaluated in continuous control benchmark\nwork, showing superior performance as well as time efficiency compared to the\nprevious methods.\n

Record transparency

Publication details

DOI
10.48550/arxiv.2012.05417
OpenAlex
W4297804711
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.