ملف الباحث
Youngjae Kim
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
2025 · arXiv (Cornell University)
LLM inference is essential for applications like text summarization, translation, and data analysis, but the high cost of GPU instances from Cloud Service Providers (CSPs) like AWS is a major burden. This paper proposes InferSave, …