Efficient Compensation of Action for Reinforcement Learning Policies in Sim2Real
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Simulation to reality (sim-to-real) transfer is a promising alternative for training behavioral policies in reinforcement learning (RL). However, many significant differences between the simulator and the real-world environment cause policies to make inconsistent actions in simulation and reality. So policies trained in simulation often perform poorly in real-world. Researchers have explored various methods to bridge this gap, including building highly realistic simulators, implementing automated domain randomization, and employing domain adaptation method. However, highly realistic simulators and automated domain randomization methods rely heavily on extensive real-world data and complex manual processes. Domain adaptation methods require the collection and annotation of high-precision real-world data, leading to low learning efficiency and restricting these methods to specific domains. To address these challenges, this paper proposes transforming the discrepancies in the transfer process into a problem of action sequence similarity. By enhancing the similarity between policy action sequences, we aim to reinforce the consistency of policy actions made in both simulation and reality. For the challenging issue of annotating real-world data, we employ a Generative Adversarial Network (GAN) framework to construct a sim-to-real consistency loss function, thus avoiding reliance on precise real-world data sampling and calibration. To avoid a large amount of real-world data sampling, we introduce Bayesian optimization to accurately and efficiently search for the optimal parameters of the compensation module. Through extensive experiments in multiple sim-to-sim scenarios as well as sim-to-real scenarios, we demonstrate that our method significantly reduces the precision and quantity requirements for real-world data sampling while maintaining high transfer performance.
Publication details
- DOI
- 10.1109/ictai62512.2024.00129
- OpenAlex
- W4406894775
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.