ملف الباحث
H. B. Ding
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
2024 · arXiv (Cornell University)
How to efficiently serve LLMs in practice has become exceptionally challenging due to their prohibitive memory and computation requirements. In this study, we investigate optimizing the KV cache, whose memory footprint poses a critical bottleneck …