ملف الباحث

H. B. Ding

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

    2024 · arXiv (Cornell University)

    How to efficiently serve LLMs in practice has become exceptionally challenging due to their prohibitive memory and computation requirements. In this study, we investigate optimizing the KV cache, whose memory footprint poses a critical bottleneck …