ملف الباحث
Mazur, Przemysław
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning
2018 · arXiv (Cornell University)
Posterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning. Randomised value functions (RVF) can be viewed as a promising approach to scaling PSRL. However, we show …