ملف الباحث
Brendan Maginnis
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
A short variational proof of equivalence between policy gradients and soft Q learning
2017 · arXiv (Cornell University)
Two main families of reinforcement learning algorithms, Q-learning and policy gradients, have recently been proven to be equivalent when using a softmax relaxation on one part, and an entropic regularization on the other. We relate …