ملف الباحث
Wencong Xiao
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Llumnix: Dynamic Scheduling for Large Language Model Serving
2024 · arXiv (Cornell University)
Inference serving for large language models (LLMs) is the key to unleashing their potential in people's daily lives. However, efficient LLM serving remains challenging today because the requests are inherently heterogeneous and unpredictable in terms …