ملف الباحث
Seohong Park
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Time Discretization-Invariant Safe Action Repetition for Policy Gradient Methods
2021 · arXiv (Cornell University)
In reinforcement learning, continuous time is often discretized by a time scale $δ$, to which the resulting performance is known to be highly sensitive. In this work, we seek to find a $δ$-invariant algorithm for …