ملف الباحث
Shiqi He
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Mordal: Automated Pretrained Model Selection for Vision Language Models
2025 · arXiv (Cornell University)
Incorporating multiple modalities into large language models (LLMs) is a powerful way to enhance their understanding of non-textual data, enabling them to perform multimodal tasks. Vision language models (VLMs) form the fastest growing category of …