AttnRoute-MoE: Attention-Prior Routing for Mixture-of-Experts Vision Transformers
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Mixture-of-Experts (MoE) Vision Transformers improve Transformer efficiency by routing each image patch token to a sparse subset of specialized feed-forward network experts. Standard learned routers produce expert collapse early in training—one or two experts absorb the majority of tokens while others go undertrained—because they discard the self-attention score matrix \(A \in \mathbb{R}^{H \times N \times N}\) before making routing decisions. This matrix encodes pairwise relational structure between tokens that, if retained, would serve as a semantically meaningful routing prior: object-center patches attend sharply and resemble the CLS token (low entropy, high CLS-similarity); boundary patches attract head disagreement (high cross-head variance); background patches attend diffusely (high entropy). We propose \(\textbf{AttnRoute-MoE},\) which constructs a structured routing prior from three properties of \(A\)—per-token attention entropy \(H_{i}\), cross-head variance \(V_{i}\), and CLS-token similarity \(C_{i}\)—and projects them onto \(E\) learned expert prototype vectors to produce a prior routing distribution \(\mathbf{P}_i \in \mathbb{R}^E\). The combined routing logit \(S_i = W_r(z_i) + \lambda_t \cdot \mathbf{P}_i\) uses a cosine annealing schedule that decays \(\lambda _{t}\) from 1 to 0 over \(T_{\mathrm{anneal}}\) steps, leaving a standard MoE with no inference overhead at the end of training. We empirically select the backbone through an attention prior diversity task (H1) comparing MAE-ViT-B/16, DeiT-B/16, DINOv1, and DINOv2-B/14. DeiT-B/16 achieves the strongest prior diversity (mean \(H\) gap \(=0.1547\) nats, \(87.6\%\) positive rate) and is used for all experiments. On Tiny ImageNet, AttnRoute-MoE reduces expert collapse (CV at epoch 1) by \(24.2\%\) in a single seed; across three random seeds the mean reduction is \(10.8\% \pm 15.4\%\) (\(t = 1.47\), \(p > 0.05\) at \(n = 3\)), with the benefit sensitive to prototype initialization. A more robust multi-seed finding is that our framework produces smoother routing trajectories (mean per-token routing variance decreased by \(14.3\%\)), yielding more stable representation learning across all experts.
Publication details
- DOI
- 10.5281/zenodo.21635491
- OpenAlex
- W7171493905
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.