ملف الباحث
Jia Yuan Yu
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Transition-based versus State-based Reward Functions for MDPs with Value-at-Risk
2016 · arXiv (Cornell University)
In reinforcement learning, the reward function on current state and action is widely used. When the objective is about the expectation of the (discounted) total reward only, it works perfectly. However, if the objective involves …