conference-paper
Open access
Leveraging Pronoun Disambiguation in Multimodal Interaction for Contextual Understanding of Voice Assistant Queries
Research footprint
At a glance
- Citations
- 1
- References
- 19
- Comments
- 0
Paper overview
Öz
Voice Assistants (VAs) are becoming an increasingly important part of our lives. However, most widespread VAs generally fail to take into account the user’s spatiotemporal context [11], leading to more descriptive and less natural dialogue. This paper introduces VOICE, an open-source multimodal VA leveraging multimodal interaction and vision-language models to allow for a more flexible and natural communication. Additionally, we present a preliminary user study to evaluate VOICE’s ability to understand queries with contextual references.
Record transparency
Publication details
- DOI
- 10.1145/3708557.3716362
- OpenAlex
- W4408550683
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.