Researcher profile
Zhu, Jianian
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
2025 · arXiv (Cornell University)
Multi-modal Large Language Models (MLLMs) serving systems commonly employ KV-cache compression to reduce memory footprint. However, existing compression methods introduce significant processing overhead and queuing delays, particularly in concurrent serving scenarios. We present \texttt{FastCache}, a …