Baoxiang Wang
3 papers in the PaperMetrix corpus
Papers by this author
-
The Gambler's Problem and Beyond
2019 · arXiv (Cornell University)
We analyze the Gambler's problem, a simple reinforcement learning problem where the gambler has the chance to double or lose the bets until the target is reached. This is an early example introduced in the …
-
Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning
2021
Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process …
-
Logarithmic Regret for Linear Markov Decision Processes with Adversarial Corruptions
2025 · Proceedings of the AAAI Conference on Artificial Intelligence
In this work, we study the logarithmic regret for reinforcement learning (RL) with linear function approximation and adversarial corruptions, in the formulation of linear Markov decision processes (MDPs). Specifically, we consider the case where there …