conference-paper

RTT- or Bandwidth-Bound? Demystifying the KV Cache Transfer in Large Language Model Serving

Research footprint

At a glance

الاستشهادات
0
المراجع
6
Comments
0
Paper overview

Abstract

Modern large language model (LLM) serving systems increasingly adopt a prefill-decode disaggregation architecture to enhance inference efficiency. While this design improves resource utilization, it introduces latency due to the transfer of key-value (KV) cache. The community has generally assumed that this latency is bandwidth-bound and can be effectively mitigated by high-speed interconnects.

Record transparency

Publication details

DOI
10.1145/3748273.3749196
OpenAlex
W4413921941
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.