article Open access

Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network

  • IEEE Transactions on Neural Networks and Learning Systems
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
27
References
52
Comments
0
Paper overview

Abstract

The deep Q-network (DQN) and return-based reinforcement learning are two promising algorithms proposed in recent years. The DQN brings advances to complex sequential decision problems, while return-based algorithms have advantages in making use of sample trajectories. In this brief, we propose a general framework to combine the DQN and most of the return-based reinforcement learning algorithms, named R-DQN. We show that the performance of the traditional DQN can be significantly improved by introducing return-based algorithms. In order to further improve the R-DQN, we design a strategy with two measurements to qualitatively measure the policy discrepancy. We conduct experiments on several representative tasks from the OpenAI Gym and Atari games. The state-of-the-art performance achieved by our method with this proposed strategy validates its effectiveness.

Record transparency

Publication details

DOI
10.1109/tnnls.2019.2948892
OpenAlex
W2809208194
Document type
article
Language
EN
Source
IEEE Transactions on Neural Networks and Learning Systems
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.