ملف الباحث
Zhiqiang Yuan
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
A Visual Leap in Clip Compositionality Reasoning Through Generation of Counterfactual Sets
2025
Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method …