ملف الباحث

Weijun Wang

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse

    2025 · arXiv (Cornell University)

    Recent advances in long-text understanding have pushed the context length of large language models (LLMs) up to one million tokens. It boosts LLMs's accuracy and reasoning capacity but causes exorbitant computational costs and unsatisfactory Time …