Researcher profile
Shuhuai Ren
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
2024 · arXiv (Cornell University)
The visual projector, which bridges the vision and language modalities and facilitates cross-modal alignment, serves as a crucial component in MLLMs. However, measuring the effectiveness of projectors in vision-language alignment remains under-explored, which currently can …
-
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
2025
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The …