conference-paper Open access

Why long model-based rollouts are no reason for bad Q-value estimates

Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Öz

This paper explores the use of model-based offline reinforcement learning with long model rollouts.While some literature criticizes this approach due to compounding errors, many practitioners have found success in real-world applications.The paper aims to demonstrate that long rollouts do not necessarily result in exponentially growing errors and can actually produce better Q-value estimates than model-free methods.These findings can potentially enhance reinforcement learning techniques.

Record transparency

Publication details

DOI
10.14428/esann/2024.es2024-80
OpenAlex
W4402365029
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.