ملف الباحث
Deyu Zhang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
2025 · arXiv (Cornell University)
Recent advances in long-text understanding have pushed the context length of large language models (LLMs) up to one million tokens. It boosts LLMs's accuracy and reasoning capacity but causes exorbitant computational costs and unsatisfactory Time …