Deep Deterministic Strategy Gradient Method Using Plot Experience Playback
At a glance
- الاستشهادات
- 0
- المراجع
- 6
- Comments
- 0
Abstract
The research on continuous control in reinforcement learning has been a hot topic in recent years. The Deep Deterministic Policy Gradient (DDPG) algorithm performs well in continuous control tasks. DDPG algorithm uses experience replay mechanism to train the network model, in order to further improve the efficiency of experience replay mechanism in the DDPG algorithm, the cumulative reward is used as the transiton classification basis, a Deep Deterministic Policy Gradient with Episodic Experience Replay (EER-DDPG) algorithm is proposed. First of all, the transitions are stored in the unit of episode, and two replay buffers are introduced respectively to classify the transitions according to the cumulative reward. Then, the quality of policy can be improved in network model training period by random samping of the episodes with large cumulative rewards. In the continuous control tasks, this algorithm is verified by experiments, and compared with DDPG algorithm, Trust Region Policy Optimization (TRPO) algorithm and Proximal Policy Optization (PPO) algorithm. The experimental results show that EER-DDPG algorithm has better performance.
Publication details
- DOI
- 10.1109/icis61260.2024.10778357
- OpenAlex
- W4405306031
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.