Researcher profile
Renrui Zhang
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
2025
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The …
-
Can World Models Benefit VLMs for World Dynamics?
2025 · arXiv (Cornell University)
Trained on internet-scale video data, generative world models are increasingly recognized as powerful world simulators that can generate consistent and plausible dynamics over structure, motion, and physics. This raises a natural question: with the advent …