Researcher profile

Zeyu Huang

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Mixture of Attention Heads: Selecting Attention Heads Per Token

    2022

    Mixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing. However, the study of MoE components mostly focused on the feedforward layer in Transformer architecture. This …