Wenjia Meng
3 papers in the PaperMetrix corpus
Papers by this author
-
Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network
2019 · IEEE Transactions on Neural Networks and Learning Systems
The deep Q-network (DQN) and return-based reinforcement learning are two promising algorithms proposed in recent years. The DQN brings advances to complex sequential decision problems, while return-based algorithms have advantages in making use of sample …
-
An Off-Policy Trust Region Policy Optimization Method With Monotonic Improvement Guarantee for Deep Reinforcement Learning
2021 · IEEE Transactions on Neural Networks and Learning Systems
In deep reinforcement learning, off-policy data help reduce on-policy interaction with the environment, and the trust region policy optimization (TRPO) method is efficient to stabilize the policy optimization procedure. In this article, we propose an …
-
MTRL-CG: Multi-Task Reinforcement Learning Method with Spectral Clustering-Based Task Grouping
2026 · Proceedings of the AAAI Conference on Artificial Intelligence
Multi-task reinforcement learning (RL) aims to enhance agent performance across multiple tasks by enabling effective knowledge transfer. However, these methods adopt a fully shared policy across all tasks without explicitly distinguishing between related and conflicting …