Zekun Li
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
2024 · arXiv (Cornell University)
Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmarks primarily focus on simple charts or other relatively straightforward figures from limited …
-
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
2024 · arXiv (Cornell University)
The rapid evolution of multimodal foundation models has led to significant advancements in cross-modal understanding and generation across diverse modalities, including text, images, audio, and video. However, these models remain susceptible to jailbreak attacks, which …
-
Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks
2025 · arXiv (Cornell University)
We study multi-turn multi-agent orchestration, where multiple large language model (LLM) agents interact over multiple turns by iteratively proposing answers or casting votes until reaching consensus. Using four LLMs (Gemini 2.5 Pro, GPT-5, Grok 4, …