Qiwen Cui
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Minimax Sample Complexity for Turn-based Stochastic Game
2021
The empirical success of multi-agent reinforcement learning is encouraging, while few theoretical guarantees have been revealed. In this work, we prove that the plug-in solver approach, probably the most natural reinforcement learning algorithm, achieves minimax …
-
Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning
2023 · arXiv (Cornell University)
Off-policy dynamic programming (DP) techniques such as $Q$-learning have proven to be important in sequential decision-making problems. In the presence of function approximation, however, these techniques often diverge due to the absence of Bellman completeness …
-
$\mathbf{(N,K)}$-Puzzle: A Cost-Efficient Testbed for Benchmarking Reinforcement Learning Algorithms in Generative Language Model
2024 · arXiv (Cornell University)
Recent advances in reinforcement learning (RL) algorithms aim to enhance the performance of language models at scale. Yet, there is a noticeable absence of a cost-effective and standardized testbed tailored to evaluating and comparing these …