ملف الباحث
Aoming Liu
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
2025 · arXiv (Cornell University)
Token reduction accelerates Multimodal Large Language Models (MLLMs) by reducing excessive tokens, but overlooks structural redundancy differences, where critical and redundant modules process identical token loads. For fine-grained computation control, we define an ``operation" as …