article

Nonparametric Bayesian Learning of Other Agents? Policies in Interactive POMDPs

  • Adaptive Agents and Multi-Agents Systems
Research footprint

At a glance

Citations
3
References
22
Comments
0
Paper overview

Öz

We consider an autonomous POMDP agent facing a multiagent environment with unknown opponents, that are modeled as finite state controllers. The agent first learns the models from (imperfectly) observed behavior, and subsequently exploits them in planning for its own optimal policy by constructing an interactive POMDP. In the learning phase, Bayesian nonparametric methods are used to sample from the posterior distribution over the infinite-dimensional space of all possible controllers, resulting in models whose size scales with the complexity of observed behavior. Experimental results show that learning improves the agent's performance, which increases with the amount of data collected during the learning phase.

Record transparency

Publication details

DOI
10.5555/2772879.2773481
OpenAlex
W2211585684
Document type
article
Language
EN
Source
Adaptive Agents and Multi-Agents Systems
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.