ملف الباحث

Jianlong Wu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

    2024 · arXiv (Cornell University)

    Large Language Models (LLMs) face significant deployment challenges due to their substantial memory requirements and the computational demands of auto-regressive text generation process. This paper addresses these challenges by focusing on the quantization of LLMs, …