ملف الباحث
Eric Kim
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
2024 · arXiv (Cornell University)
Efficient inference in large language models (LLMs) has become a critical focus as their scale and complexity grow. Traditional autoregressive decoding, while effective, suffers from computational inefficiencies due to its sequential token generation process. Speculative …