An N-step Look Ahead Algorithm Using Mixed (On and Off) Policy Reinforcement Learning
At a glance
- Citations
- 2
- References
- 12
- Comments
- 0
Abstract
In this paper, a new algorithm for reinforcement learning is proposed. Q-learning and SARSA are one of the most famous algorithms used in Reinforcement learning. Q-learning gives the most optimal path whereas SARSA gives the safest path. This paper proposes an algorithm that allows users to flexibly chose between safer or faster path in any given environment by altering the hyper parameter n. The approach taken to achieve this is N-step look ahead. In this the agent will update the initial states value after it has taken n steps using on- policy reinforcement learning and in adjacent step it applies off- policy RL. Using this technique, the algorithm overcomes the overestimation of value function prevalent in Q-learning as it emphasizes more on on-policy updates. This algorithm allows the user alter the dependency on bootstrapping and sampling in updating the value function.
Publication details
- DOI
- 10.1109/iciss49785.2020.9315959
- OpenAlex
- W3126245445
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.