ملف الباحث
Mudit Verma
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Explanation Augmented Feedback in Human-in-the-Loop Reinforcement Learning
2020 · arXiv (Cornell University)
Human-in-the-loop Reinforcement Learning (HRL) aims to integrate human guidance with Reinforcement Learning (RL) algorithms to improve sample efficiency and performance. A common type of human guidance in HRL is binary evaluative good or bad feedback …
-
Computing Policies That Account For The Effects Of Human Agent Uncertainty During Execution In Markov Decision Processes
2021 · arXiv (Cornell University)
When humans are given a policy to execute, there can be policy execution errors and deviations in policy if there is uncertainty in identifying a state. This can happen due to the human agent's cognitive …