ملف الباحث
Fuwen Tian
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language Models
2026 · Proceedings of the AAAI Conference on Artificial Intelligence
Large Language Models (LLMs) have revolutionized intelligent interactions, enabling mobile applications such as personal assistants on edge devices for local execution. Speculative decoding (SD) has emerged as a promising paradigm to accelerate LLM inference without …