Researcher profile

Guangming Cui

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language Models

    2026 · Proceedings of the AAAI Conference on Artificial Intelligence

    Large Language Models (LLMs) have revolutionized intelligent interactions, enabling mobile applications such as personal assistants on edge devices for local execution. Speculative decoding (SD) has emerged as a promising paradigm to accelerate LLM inference without …