conference-paper

Policy improvement in dynamic programming

  • 2nd International Conference on Artificial Intelligence, Automation, and High-Performance Computing (AIAHPC 2022)
Research footprint

At a glance

Citations
0
References
9
Comments
0
Paper overview

Öz

Policy improvement has a long history and is the essential element in dynamic programming. The general categories of policy improvement can be divided into four aspects including: heuristic methods, approximation methods, sampling methods and numerical improvement. Paralleling with the classic policy improvement methods, several variant tools are also introduced including Lambda Policy, Path Integral, High Confidence in Policy Improvement and Finite Sample Analysis for SARSA in linear function. Moreover, the introductions of those policy improvement methods, evaluation, and comparison between them are illustrated in this paper. There are totally three perspectives where this paper dissects the evaluation from training speed, sampling efficiency and methods ability.

Record transparency

Publication details

DOI
10.1117/12.2641811
OpenAlex
W4308702803
Document type
conference-paper
Language
EN
Source
2nd International Conference on Artificial Intelligence, Automation, and High-Performance Computing (AIAHPC 2022)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.