ملف الباحث

Hossein Entezari Zarch

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding

    2025 · arXiv (Cornell University)

    Speculative Decoding (SD) is a widely used approach to accelerate the inference of large language models (LLMs) without reducing generation quality. It operates by first using a compact model to draft multiple tokens efficiently, followed …