ملف الباحث

Yao, Feiyu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

    2025 · arXiv (Cornell University)

    As large language models (LLMs) continue to support increasingly longer contexts, the memory demand for key-value (KV) caches during decoding grows rapidly, becoming a critical bottleneck in both GPU memory capacity and PCIe bandwidth. Sparse …