Researcher profile
Rongzhi Zhang
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
2024 · arXiv (Cornell University)
The Key-Value (KV) cache is a crucial component in serving transformer-based autoregressive large language models (LLMs), enabling faster inference by storing previously computed KV vectors. However, its memory consumption scales linearly with sequence length and …