Researcher profile
Gang Pan
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network
2019 · IEEE Transactions on Neural Networks and Learning Systems
The deep Q-network (DQN) and return-based reinforcement learning are two promising algorithms proposed in recent years. The DQN brings advances to complex sequential decision problems, while return-based algorithms have advantages in making use of sample …
-
An Off-Policy Trust Region Policy Optimization Method With Monotonic Improvement Guarantee for Deep Reinforcement Learning
2021 · IEEE Transactions on Neural Networks and Learning Systems
In deep reinforcement learning, off-policy data help reduce on-policy interaction with the environment, and the trust region policy optimization (TRPO) method is efficient to stabilize the policy optimization procedure. In this article, we propose an …