conference-paper
وصول مفتوح
Why long model-based rollouts are no reason for bad Q-value estimates
Research footprint
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Paper overview
Abstract
This paper explores the use of model-based offline reinforcement learning with long model rollouts.While some literature criticizes this approach due to compounding errors, many practitioners have found success in real-world applications.The paper aims to demonstrate that long rollouts do not necessarily result in exponentially growing errors and can actually produce better Q-value estimates than model-free methods.These findings can potentially enhance reinforcement learning techniques.
Record transparency
Publication details
- DOI
- 10.14428/esann/2024.es2024-80
- OpenAlex
- W4402365029
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.