conference-paper

Improving Open-Ended Referring Expression Comprehension via Dual-Language Constraints

Research footprint

At a glance

Citations
0
References
35
Comments
0
Paper overview

Abstract

Open-ended referring expression comprehension focuses on locating the text query within an image via scene knowledge, requiring complex reasoning across the triplet of the image, scene knowledge, and the text query. However, most existing methods struggle to integrate scene knowledge while performing a single-object prediction. To address this issue, we propose VG-DLC, a model that progressively uses scene knowledge and open-ended referring expressions as constraints for reasoning and grounding. Specifically, we first align scene knowledge with the referring expression and the image in sequence, which supports effective relational reasoning and implicitly constrains the open-ended content to corresponding image regions. Next, we explicitly align the referring expression with the image and employ it as a query vector, constraining it to a unique region to achieve a single end-to-end prediction. Extensive experiments on the SK-VG dataset validate the effectiveness of our method, demonstrating that VG-DLC outperforms existing end-to-end approaches.

Record transparency

Publication details

DOI
10.1109/icassp49660.2025.10888443
OpenAlex
W4408354313
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.