ملف الباحث

Jia Yuan Yu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Transition-based versus State-based Reward Functions for MDPs with Value-at-Risk

    2016 · arXiv (Cornell University)

    In reinforcement learning, the reward function on current state and action is widely used. When the objective is about the expectation of the (discounted) total reward only, it works perfectly. However, if the objective involves …