Researcher profile
Z. Zhang
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
RTT- or Bandwidth-Bound? Demystifying the KV Cache Transfer in Large Language Model Serving
2025
Modern large language model (LLM) serving systems increasingly adopt a prefill-decode disaggregation architecture to enhance inference efficiency. While this design improves resource utilization, it introduces latency due to the transfer of key-value (KV) cache. The …