ملف الباحث
Bozhi Luan
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
2024 · arXiv (Cornell University)
The advent of Large Multimodal Models (LMMs) has sparked a surge in research aimed at harnessing their remarkable reasoning abilities. However, for understanding text-rich images, challenges persist in fully leveraging the potential of LMMs, and …