conference-paper

Deep Deterministic Strategy Gradient Method Using Plot Experience Playback

Research footprint

At a glance

الاستشهادات
0
المراجع
6
Comments
0
Paper overview

Abstract

The research on continuous control in reinforcement learning has been a hot topic in recent years. The Deep Deterministic Policy Gradient (DDPG) algorithm performs well in continuous control tasks. DDPG algorithm uses experience replay mechanism to train the network model, in order to further improve the efficiency of experience replay mechanism in the DDPG algorithm, the cumulative reward is used as the transiton classification basis, a Deep Deterministic Policy Gradient with Episodic Experience Replay (EER-DDPG) algorithm is proposed. First of all, the transitions are stored in the unit of episode, and two replay buffers are introduced respectively to classify the transitions according to the cumulative reward. Then, the quality of policy can be improved in network model training period by random samping of the episodes with large cumulative rewards. In the continuous control tasks, this algorithm is verified by experiments, and compared with DDPG algorithm, Trust Region Policy Optimization (TRPO) algorithm and Proximal Policy Optization (PPO) algorithm. The experimental results show that EER-DDPG algorithm has better performance.

Record transparency

Publication details

DOI
10.1109/icis61260.2024.10778357
OpenAlex
W4405306031
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.