ملف الباحث

Jiguang Wan

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference

    2025 · arXiv (Cornell University)

    Recently, large language models (LLMs) have been able to handle longer and longer contexts. However, a context that is too long may cause intolerant inference latency and GPU memory usage. Existing methods propose mixed-precision quantization …