preprint وصول مفتوح

Agents that Listen: High-Throughput Reinforcement Learning with Multiple\n Sensory Systems

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
0
المراجع
0
Comments
0
Paper overview

Abstract

Humans and other intelligent animals evolved highly sophisticated perception\nsystems that combine multiple sensory modalities. On the other hand,\nstate-of-the-art artificial agents rely mostly on visual inputs or structured\nlow-dimensional observations provided by instrumented environments. Learning to\nact based on combined visual and auditory inputs is still a new topic of\nresearch that has not been explored beyond simple scenarios. To facilitate\nprogress in this area we introduce a new version of VizDoom simulator to create\na highly efficient learning environment that provides raw audio observations.\nWe study the performance of different model architectures in a series of tasks\nthat require the agent to recognize sounds and execute instructions given in\nnatural language. Finally, we train our agent to play the full game of Doom and\nfind that it can consistently defeat a traditional vision-based adversary. We\nare currently in the process of merging the augmented simulator with the main\nViZDoom code repository. Video demonstrations and experiment code can be found\nat https://sites.google.com/view/sound-rl.\n

Record transparency

Publication details

DOI
10.48550/arxiv.2107.02195
OpenAlex
W4287100169
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.