ملف الباحث
Zhu, Jianian
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
2025 · arXiv (Cornell University)
Multi-modal Large Language Models (MLLMs) serving systems commonly employ KV-cache compression to reduce memory footprint. However, existing compression methods introduce significant processing overhead and queuing delays, particularly in concurrent serving scenarios. We present \texttt{FastCache}, a …