ملف الباحث
Linli Yao
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
2024 · arXiv (Cornell University)
The visual projector, which bridges the vision and language modalities and facilitates cross-modal alignment, serves as a crucial component in MLLMs. However, measuring the effectiveness of projectors in vision-language alignment remains under-explored, which currently can …