ملف الباحث
Andrés Marafioti
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Building and better understanding vision-language models: insights and future directions
2024 · arXiv (Cornell University)
The field of vision-language models (VLMs), which take images and texts as inputs and output texts, is rapidly evolving and has yet to reach consensus on several key aspects of the development pipeline, including data, …