Researcher profile

Shuhuai Ren

2 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

    2024 · arXiv (Cornell University)

    The visual projector, which bridges the vision and language modalities and facilitates cross-modal alignment, serves as a crucial component in MLLMs. However, measuring the effectiveness of projectors in vision-language alignment remains under-explored, which currently can …

  2. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    2025

    In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The …