ملف الباحث

Anirudh Atmakuru

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Does quantization affect models' performance on long-context tasks?

    2025 · arXiv (Cornell University)

    Large language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency. Quantization can mitigate these costs, but may degrade performance. In this work, we …