ملف الباحث
Stephan Alaniz
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
2024 · arXiv (Cornell University)
Vision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly on foreground objects, neglecting other crucial …