TEAC: Intergrating Trust Region and Max Entropy Actor Critic for Continuous Control
At a glance
- الاستشهادات
- 3
- المراجع
- 0
- Comments
- 0
Abstract
Trust region methods and maximum entropy methods are two state-of-the-art branches used in reinforcement learning (RL) for the benefits of stability and exploration in continuous environments, respectively. This paper proposes to integrate both branches in a unified framework, thus benefiting from both sides. We first transform the original RL objective to a constraint optimization problem and then proposes trust entropy actor-critic (TEAC), an off-policy algorithm to learn stable and sufficiently explored policies for continuous states and actions. TEAC trains the critic by minimizing the refined Bellman error and updates the actor by minimizing KL-divergence loss derived from the closed-form solution to the Lagrangian. We prove that the policy evaluation and policy improvement in TEAC is guaranteed to converge. We compare TEAC with 4 state-of-the-art solutions on 6 tasks in the MuJoCo environment. The results show that TEAC outperforms state-of-the-art solutions in terms of efficiency and effectiveness.
Publication details
- OpenAlex
- W3133371913
- Document type
- article
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.