conference-paper
RTT- or Bandwidth-Bound? Demystifying the KV Cache Transfer in Large Language Model Serving
Research footprint
At a glance
- Citations
- 0
- References
- 6
- Comments
- 0
Paper overview
Öz
Modern large language model (LLM) serving systems increasingly adopt a prefill-decode disaggregation architecture to enhance inference efficiency. While this design improves resource utilization, it introduces latency due to the transfer of key-value (KV) cache. The community has generally assumed that this latency is bandwidth-bound and can be effectively mitigated by high-speed interconnects.
Record transparency
Publication details
- DOI
- 10.1145/3748273.3749196
- OpenAlex
- W4413921941
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.