Policy improvement in dynamic programming
At a glance
- Citations
- 0
- References
- 9
- Comments
- 0
Öz
Policy improvement has a long history and is the essential element in dynamic programming. The general categories of policy improvement can be divided into four aspects including: heuristic methods, approximation methods, sampling methods and numerical improvement. Paralleling with the classic policy improvement methods, several variant tools are also introduced including Lambda Policy, Path Integral, High Confidence in Policy Improvement and Finite Sample Analysis for SARSA in linear function. Moreover, the introductions of those policy improvement methods, evaluation, and comparison between them are illustrated in this paper. There are totally three perspectives where this paper dissects the evaluation from training speed, sampling efficiency and methods ability.
Publication details
- DOI
- 10.1117/12.2641811
- OpenAlex
- W4308702803
- Document type
- conference-paper
- Language
- EN
- Source
- 2nd International Conference on Artificial Intelligence, Automation, and High-Performance Computing (AIAHPC 2022)
- Last metadata update
Comments
Oturum Açın to join the discussion.