Researcher profile
Guangming Cui
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language Models
2026 · Proceedings of the AAAI Conference on Artificial Intelligence
Large Language Models (LLMs) have revolutionized intelligent interactions, enabling mobile applications such as personal assistants on edge devices for local execution. Speculative decoding (SD) has emerged as a promising paradigm to accelerate LLM inference without …