conference-paper

An N-step Look Ahead Algorithm Using Mixed (On and Off) Policy Reinforcement Learning

Research footprint

At a glance

Citations
2
References
12
Comments
0
Paper overview

Öz

In this paper, a new algorithm for reinforcement learning is proposed. Q-learning and SARSA are one of the most famous algorithms used in Reinforcement learning. Q-learning gives the most optimal path whereas SARSA gives the safest path. This paper proposes an algorithm that allows users to flexibly chose between safer or faster path in any given environment by altering the hyper parameter n. The approach taken to achieve this is N-step look ahead. In this the agent will update the initial states value after it has taken n steps using on- policy reinforcement learning and in adjacent step it applies off- policy RL. Using this technique, the algorithm overcomes the overestimation of value function prevalent in Q-learning as it emphasizes more on on-policy updates. This algorithm allows the user alter the dependency on bootstrapping and sampling in updating the value function.

Record transparency

Publication details

DOI
10.1109/iciss49785.2020.9315959
OpenAlex
W3126245445
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.