conference-paper وصول مفتوح

Why long model-based rollouts are no reason for bad Q-value estimates

Research footprint

At a glance

الاستشهادات
0
المراجع
0
Comments
0
Paper overview

Abstract

This paper explores the use of model-based offline reinforcement learning with long model rollouts.While some literature criticizes this approach due to compounding errors, many practitioners have found success in real-world applications.The paper aims to demonstrate that long rollouts do not necessarily result in exponentially growing errors and can actually produce better Q-value estimates than model-free methods.These findings can potentially enhance reinforcement learning techniques.

Record transparency

Publication details

DOI
10.14428/esann/2024.es2024-80
OpenAlex
W4402365029
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.