preprint Open access

AttnRoute-MoE: Attention-Prior Routing for Mixture-of-Experts Vision Transformers

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Mixture-of-Experts (MoE) Vision Transformers improve Transformer efficiency by routing each image patch token to a sparse subset of specialized feed-forward network experts. Standard learned routers produce expert collapse early in training—one or two experts absorb the majority of tokens while others go undertrained—because they discard the self-attention score matrix \(A \in \mathbb{R}^{H \times N \times N}\) before making routing decisions. This matrix encodes pairwise relational structure between tokens that, if retained, would serve as a semantically meaningful routing prior: object-center patches attend sharply and resemble the CLS token (low entropy, high CLS-similarity); boundary patches attract head disagreement (high cross-head variance); background patches attend diffusely (high entropy). We propose \(\textbf{AttnRoute-MoE},\) which constructs a structured routing prior from three properties of \(A\)—per-token attention entropy \(H_{i}\), cross-head variance \(V_{i}\), and CLS-token similarity \(C_{i}\)—and projects them onto \(E\) learned expert prototype vectors to produce a prior routing distribution \(\mathbf{P}_i \in \mathbb{R}^E\). The combined routing logit \(S_i = W_r(z_i) + \lambda_t \cdot \mathbf{P}_i\) uses a cosine annealing schedule that decays \(\lambda _{t}\) from 1 to 0 over \(T_{\mathrm{anneal}}\) steps, leaving a standard MoE with no inference overhead at the end of training. We empirically select the backbone through an attention prior diversity task (H1) comparing MAE-ViT-B/16, DeiT-B/16, DINOv1, and DINOv2-B/14. DeiT-B/16 achieves the strongest prior diversity (mean \(H\) gap \(=0.1547\) nats, \(87.6\%\) positive rate) and is used for all experiments. On Tiny ImageNet, AttnRoute-MoE reduces expert collapse (CV at epoch 1) by \(24.2\%\) in a single seed; across three random seeds the mean reduction is \(10.8\% \pm 15.4\%\) (\(t = 1.47\), \(p > 0.05\) at \(n = 3\)), with the benefit sensitive to prototype initialization. A more robust multi-seed finding is that our framework produces smoother routing trajectories (mean per-token routing variance decreased by \(14.3\%\)), yielding more stable representation learning across all experts.

Record transparency

Publication details

DOI
10.5281/zenodo.21635491
OpenAlex
W7171493905
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.