Researcher profile
Zhen Wu
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
2025 · arXiv (Cornell University)
Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires loading all expert parameters, …