ملف الباحث

Wencong Xiao

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Llumnix: Dynamic Scheduling for Large Language Model Serving

    2024 · arXiv (Cornell University)

    Inference serving for large language models (LLMs) is the key to unleashing their potential in people's daily lives. However, efficient LLM serving remains challenging today because the requests are inherently heterogeneous and unpredictable in terms …