Noise Accumulation and Rank Collapse in Dense Self-Attention: DSALT
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Large language models based on the Transformer decoder architecture perform multi-head self-attention over all previous tokens in the context window. We argue that this dense attention mechanism introduces a systematic form of noise: every token, regardless of semantic relevance, contributes a strictly positive weight to every other token's representation via the softmax operation. This noise accumulates across attention heads and layers, progressively corrupting token representations. We show that this accumulation accelerates the rank collapse phenomenon established by Dong et al., in which self-attention networks converge doubly exponentially to a rank-1 matrix with depth. We conjecture that this mechanism is a structural cause of hallucinations in large language models, consistent with empirical evidence on long-context degradation. To address this, we propose Dynamic Sparse Attention with Landmark Tokens (DSALT), a mechanism that replaces dense attention with an adaptive local window augmented by a small set of globally informative tokens, reducing noise at its source while preserving essential long-range dependencies. We validate these claims empirically: in a controlled language modeling experiment with identical model size and training budget, DSALT achieves a 9% reduction in validation perplexity (172.62 -> 156.78) and a 13% reduction in sliding-window perplexity over longer contexts (152.86 -> 132.97), with the gap widening on longer sequences, precisely as the cumulative noise hypothesis predicts. A 15% reduction in directly measured noise norm further confirms the mechanism, with no overhead in per-step computational cost.
Publication details
- DOI
- 10.5281/zenodo.19384947
- OpenAlex
- W7148327613
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.