ملف الباحث
Olivier Bachem
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
A functional mirror ascent view of policy gradient methods with function approximation.
2021 · arXiv (Cornell University)
We use functional mirror ascent to propose a general framework (referred to as FMA-PG) for designing policy gradient methods. The functional perspective distinguishes between a policy's functional representation (what are its sufficient statistics) and its …
-
The Role of Pretrained Representations for the OOD Generalization of Reinforcement Learning Agents
2021 · arXiv (Cornell University)
Building sample-efficient agents that generalize out-of-distribution (OOD) in real-world settings remains a fundamental unsolved problem on the path towards achieving higher-level cognition. One particularly promising approach is to begin with low-dimensional, pretrained representations of our …