ملف الباحث
Johan Ferret
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Adversarially Guided Actor-Critic
2021 · arXiv (Cornell University)
Despite definite success in deep reinforcement learning problems,\nactor-critic algorithms are still confronted with sample inefficiency in\ncomplex environments, particularly in tasks where efficient exploration is a\nbottleneck. These methods consider a policy (the actor) and a value …
-
Self-Imitation Advantage Learning
2021
Self-imitation learning is a Reinforcement Learning (RL) method that encourages actions whose returns were higher than expected, which helps in hard exploration and sparse reward problems. It was shown to improve the performance of on-policy …