preprint Open access

Semi-On-Policy Training for Sample Efficient Multi-Agent Policy Gradients

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
1
References
35
Comments
0
Paper overview

Öz

Policy gradient methods are an attractive approach to multi-agent reinforcement learning problems due to their convergence properties and robustness in partially observable scenarios. However, there is a significant performance gap between state-of-the-art policy gradient and value-based methods on the popular StarCraft Multi-Agent Challenge (SMAC) benchmark. In this paper, we introduce semi-on-policy (SOP) training as an effective and computationally efficient way to address the sample inefficiency of on-policy policy gradient methods. We enhance two state-of-the-art policy gradient algorithms with SOP training, demonstrating significant performance improvements. Furthermore, we show that our methods perform as well or better than state-of-the-art value-based methods on a variety of SMAC tasks.

Record transparency

Publication details

DOI
10.48550/arxiv.2104.13446
OpenAlex
W3158227719
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.