ملف الباحث
Niao He
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Provably Learning Nash Policies in Constrained Markov Potential Games
2023 · arXiv (Cornell University)
Multi-agent reinforcement learning (MARL) addresses sequential decision-making problems with multiple agents, where each agent optimizes its own objective. In many real-world instances, the agents may not only want to optimize their objectives, but also ensure …
-
Provably Convergent Policy Optimization via Metric-aware Trust Region Methods
2023 · arXiv (Cornell University)
Trust-region methods based on Kullback-Leibler divergence are pervasively used to stabilize policy optimization in reinforcement learning. In this paper, we exploit more flexible metrics and examine two natural extensions of policy optimization with Wasserstein and …