Deep Reinforcement Learning algorithms learn important classes of repeated games optimally—Theoretical and empirical analysis
At a glance
- Citations
- 0
- References
- 8
- Comments
- 0
Abstract
This paper evaluates two prominent Deep Reinforcement Learning algorithms, Deep Q-Learning and Twin Delayed Deep Deterministic Policy Gradient, by comparing their learned policies against analytically derived optimal policies in specific game-theoretic models. We employ repeated, symmetric and generic 2 × 2 games and a repeated Cournot duopoly, both featuring a single Deep Reinforcement Learning agent playing against a myopic Best Response Player. Our empirical results demonstrate that the agents reliably learn to replicate these optimal policies. This success suggests the potential of Reinforcement Learning as a useful tool for game-theoretic analysis in more complex environments. To provide a theoretical anchor for these findings, we present a proof of subsumption showing that Deep Q-Learning, under a specific parametrization, inherits the formal convergence guarantees of classical Q-Learning in finite Markov Decision Processes. • Performance of Deep Reinforcement Learning agents is benchmarked against analytical optima in game-theoretic models. • Optimal agent behavior is shown to be underpinned by correctly learned Q-functions, aligning with theoretical reward evaluations. • A proof of subsumption establishes that Deep Q-Learning inherits the formal convergence guarantees of classical Q-Learning under certain conditions.
Publication details
- DOI
- 10.1016/j.fraope.2026.100503
- OpenAlex
- W7123526514
- Document type
- article
- Language
- EN
- Source
- Franklin Open
- Last metadata update
Comments
Log in to join the discussion.