Active Imitation Learning from Multiple Non-Deterministic Teachers:\n Formulation, Challenges, and Algorithms
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
We formulate the problem of learning to imitate multiple, non-deterministic\nteachers with minimal interaction cost. Rather than learning a specific policy\nas in standard imitation learning, the goal in this problem is to learn a\ndistribution over a policy space. We first present a general framework that\nefficiently models and estimates such a distribution by learning continuous\nrepresentations of the teacher policies. Next, we develop Active\nPerformance-Based Imitation Learning (APIL), an active learning algorithm for\nreducing the learner-teacher interaction cost in this framework. By making\nquery decisions based on predictions of future progress, our algorithm avoids\nthe pitfalls of traditional uncertainty-based approaches in the face of teacher\nbehavioral uncertainty. Results on both toy and photo-realistic navigation\ntasks show that APIL significantly reduces the numbers of interactions with\nteachers without compromising on performance. Moreover, it is robust to various\ndegrees of teacher behavioral uncertainty.\n
Publication details
- DOI
- 10.48550/arxiv.2006.07777
- OpenAlex
- W4287758112
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.