Adversarially Guided Actor-Critic
At a glance
- Citations
- 3
- References
- 53
- Comments
- 0
Abstract
Despite definite success in deep reinforcement learning problems,\nactor-critic algorithms are still confronted with sample inefficiency in\ncomplex environments, particularly in tasks where efficient exploration is a\nbottleneck. These methods consider a policy (the actor) and a value function\n(the critic) whose respective losses are built using different motivations and\napproaches. This paper introduces a third protagonist: the adversary. While the\nadversary mimics the actor by minimizing the KL-divergence between their\nrespective action distributions, the actor, in addition to learning to solve\nthe task, tries to differentiate itself from the adversary predictions. This\nnovel objective stimulates the actor to follow strategies that could not have\nbeen correctly predicted from previous trajectories, making its behavior\ninnovative in tasks where the reward is extremely rare. Our experimental\nanalysis shows that the resulting Adversarially Guided Actor-Critic (AGAC)\nalgorithm leads to more exhaustive exploration. Notably, AGAC outperforms\ncurrent state-of-the-art methods on a set of various hard-exploration and\nprocedurally-generated tasks.\n
Publication details
- DOI
- 10.48550/arxiv.2102.04376
- OpenAlex
- W3123169626
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.