Researcher profile
Pierre H. Richemond
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
A short variational proof of equivalence between policy gradients and soft Q learning
2017 · arXiv (Cornell University)
Two main families of reinforcement learning algorithms, Q-learning and policy gradients, have recently been proven to be equivalent when using a softmax relaxation on one part, and an entropic regularization on the other. We relate …
-
Sample-Efficient Reinforcement Learning with Maximum Entropy Mellowmax Episodic Control
2019 · arXiv (Cornell University)
Deep networks have enabled reinforcement learning to scale to more complex and challenging domains, but these methods typically require large quantities of training data. An alternative is to use sample-efficient episodic control methods: neuro-inspired algorithms …