Researcher profile
Zeyu Huang
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Mixture of Attention Heads: Selecting Attention Heads Per Token
2022
Mixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing. However, the study of MoE components mostly focused on the feedforward layer in Transformer architecture. This …