ملف الباحث
Xinhan Di
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
2024 · arXiv (Cornell University)
There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded objects for visual-language multimodal models through …