article Open access

Deep Reinforcement Learning algorithms learn important classes of repeated games optimally—Theoretical and empirical analysis

  • Franklin Open
  • Elsevier BV
Research footprint

At a glance

Citations
0
References
8
Comments
0
Paper overview

Abstract

This paper evaluates two prominent Deep Reinforcement Learning algorithms, Deep Q-Learning and Twin Delayed Deep Deterministic Policy Gradient, by comparing their learned policies against analytically derived optimal policies in specific game-theoretic models. We employ repeated, symmetric and generic 2 × 2 games and a repeated Cournot duopoly, both featuring a single Deep Reinforcement Learning agent playing against a myopic Best Response Player. Our empirical results demonstrate that the agents reliably learn to replicate these optimal policies. This success suggests the potential of Reinforcement Learning as a useful tool for game-theoretic analysis in more complex environments. To provide a theoretical anchor for these findings, we present a proof of subsumption showing that Deep Q-Learning, under a specific parametrization, inherits the formal convergence guarantees of classical Q-Learning in finite Markov Decision Processes. • Performance of Deep Reinforcement Learning agents is benchmarked against analytical optima in game-theoretic models. • Optimal agent behavior is shown to be underpinned by correctly learned Q-functions, aligning with theoretical reward evaluations. • A proof of subsumption establishes that Deep Q-Learning inherits the formal convergence guarantees of classical Q-Learning under certain conditions.

Record transparency

Publication details

DOI
10.1016/j.fraope.2026.100503
OpenAlex
W7123526514
Document type
article
Language
EN
Source
Franklin Open
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.