Researcher profile
Hossein Entezari Zarch
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
2025 · arXiv (Cornell University)
Speculative Decoding (SD) is a widely used approach to accelerate the inference of large language models (LLMs) without reducing generation quality. It operates by first using a compact model to draft multiple tokens efficiently, followed …