ملف الباحث
Sepehr Lavasani
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
2025 · ArXiv.org
Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing training-free methods typically estimate feedforward network (FFN) neuron importance from the input prompt alone. We show this prompt-only …