conference-paper

RTT- or Bandwidth-Bound? Demystifying the KV Cache Transfer in Large Language Model Serving

Research footprint

At a glance

Citations
0
References
6
Comments
0
Paper overview

Abstract

Modern large language model (LLM) serving systems increasingly adopt a prefill-decode disaggregation architecture to enhance inference efficiency. While this design improves resource utilization, it introduces latency due to the transfer of key-value (KV) cache. The community has generally assumed that this latency is bandwidth-bound and can be effectively mitigated by high-speed interconnects.

Record transparency

Publication details

DOI
10.1145/3748273.3749196
OpenAlex
W4413921941
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.