preprint Open access

ZeroMass Architecture: Mitigating the "Memory Wall" in LLM Inference via Asynchronous Decompression and Semantic Proxies

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

ZeroMass Architecture proposes a theoretical systems architecture for mitigating the Memory Wall in Large Language Model (LLM) inference on commodity hardware. Rather than treating model parameters as a static asset permanently resident in GPU memory, ZeroMass reformulates inference as a dynamic asynchronous streaming process. The architecture combines low-rank matrix decomposition (SVD), asynchronous memory pipelines, semantic context reduction through proxy graphs, and dynamic KV-cache recompression to reduce memory pressure while preserving computational efficiency. This work presents the mathematical framework, architectural design, and theoretical analysis supporting the proposed approach. It does not claim experimental validation; implementation of a production-grade runtime and empirical benchmarking are left for future work. The paper is intended as an open research contribution and is released under the Creative Commons Attribution 4.0 (CC BY 4.0) license to encourage discussion, experimentation, and further development by the research community.

Record transparency

Publication details

DOI
10.5281/zenodo.21267011
OpenAlex
W7167725072
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.