conference-paper

IoU-CLIP: IoU-Aware Language-Image Model Tuning for Open Vocabulary Object Detection

Research footprint

At a glance

Citations
0
References
28
Comments
0
Paper overview

Abstract

Open vocabulary object detection (OVD), which detects novel categories through detectors trained on base categories, has achieved remarkable advancement attributable to large-scale vision-language models, such as CLIP. The prior OVD works mainly focused on improving the classification accuracy of proposals, ignoring the ability of localization for novel categories. In this work, we propose IoU-aware language-image model tuning (IoU-CLIP) for open vocabulary object detection. Specifically, we construct a region image dataset with different IoU and adopt IoU values as labels to fine-tune the CLIP model to learn IoU-aware and class-agnostic semantic prompts and visual embeddings. The fine-tuned IoU-CLIP can predict IoU scores for proposals, which interact with classification scores. Meanwhile, IoU-aware and class-agnostic visual embeddings are utilized for box regression to enhance the generalization of the localization capability. We evaluate our method on the COCO and LVIS OVD benchmarks, outperforming the baseline (RegionCLIP) by 5.5% AP50and 5.8% AP on novel categories, respectively, achieving state-of-the-art performance.

Record transparency

Publication details

DOI
10.1109/vcip63160.2024.10849818
OpenAlex
W4406860151
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.