ملف الباحث

Cody Hao Yu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Efficient Memory Management for Large Language Model Serving with PagedAttention

    2023

    High throughput serving of large language models (LLMs) requires batching sufficiently many requests at a time. However, existing systems struggle because the key-value cache (KV cache) memory for each request is huge and grows and …