Learning and Planning for Time-Varying MDPs Using Maximum Likelihood\n Estimation
At a glance
- Citations
- 1
- References
- 0
- Comments
- 0
Abstract
This paper proposes a formal approach to online learning and planning for\nagents operating in a priori unknown, time-varying environments. The proposed\nmethod computes the maximally likely model of the environment, given the\nobservations about the environment made by an agent earlier in the system run\nand assuming knowledge of a bound on the maximal rate of change of system\ndynamics. Such an approach generalizes the estimation method commonly used in\nlearning algorithms for unknown Markov decision processes with time-invariant\ntransition probabilities, but is also able to quickly and correctly identify\nthe system dynamics following a change. Based on the proposed method, we\ngeneralize the exploration bonuses used in learning for time-invariant Markov\ndecision processes by introducing a notion of uncertainty in a learned\ntime-varying model, and develop a control policy for time-varying Markov\ndecision processes based on the exploitation and exploration trade-off. We\ndemonstrate the proposed methods on four numerical examples: a patrolling task\nwith a change in system dynamics, a two-state MDP with periodically changing\noutcomes of actions, a wind flow estimation task, and a multi-armed bandit\nproblem with periodically changing probabilities of different rewards.\n
Publication details
- DOI
- 10.48550/arxiv.1911.12976
- OpenAlex
- W4288009418
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.