An Efficient Asynchronous Method for Integrating Evolutionary and\n Gradient-based Policy Search
At a glance
- Citations
- 3
- References
- 0
- Comments
- 0
Abstract
Deep reinforcement learning (DRL) algorithms and evolution strategies (ES)\nhave been applied to various tasks, showing excellent performances. These have\nthe opposite properties, with DRL having good sample efficiency and poor\nstability, while ES being vice versa. Recently, there have been attempts to\ncombine these algorithms, but these methods fully rely on synchronous update\nscheme, making it not ideal to maximize the benefits of the parallelism in ES.\nTo solve this challenge, asynchronous update scheme was introduced, which is\ncapable of good time-efficiency and diverse policy exploration. In this paper,\nwe introduce an Asynchronous Evolution Strategy-Reinforcement Learning (AES-RL)\nthat maximizes the parallel efficiency of ES and integrates it with policy\ngradient methods. Specifically, we propose 1) a novel framework to merge ES and\nDRL asynchronously and 2) various asynchronous update methods that can take all\nadvantages of asynchronism, ES, and DRL, which are exploration and time\nefficiency, stability, and sample efficiency, respectively. The proposed\nframework and update methods are evaluated in continuous control benchmark\nwork, showing superior performance as well as time efficiency compared to the\nprevious methods.\n
Publication details
- DOI
- 10.48550/arxiv.2012.05417
- OpenAlex
- W4297804711
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.