Researcher profile

Kunlin Zhang

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs

    2025 · ArXiv.org

    Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing training-free methods typically estimate feedforward network (FFN) neuron importance from the input prompt alone. We show this prompt-only …