Regime-Specific Pressure Laws for Dropout Decay in Streaming Language Model Training
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
This record archives a preprint manuscript and reproducibility snapshot for an empirical study of dropout decay in streaming language-model training. The paper studies whether dropout should be treated as a regime-specific pressure law rather than a fixed hyperparameter. Static dropout calibration cells are used to fit interaction laws over model parameter count, currently available unique tokens, and cumulative sampled training tokens. Frozen coefficient-derived schedules are then validated against fixed-dropout baselines across OpenWebText10K, TinyStories, and WikiText-103. The archive includes the paper PDF, LaTeX source, code, scripts, tests, saved metrics/results, coefficient files, cached tokenizers/token arrays, local source corpora, and reproducibility documentation. The codebase is derived from Andrej Karpathy's nanochat architecture and retains MIT attribution. Repository: https://huggingface.co/cuber12/dropout-decayGit tag: v1.1-preprint https://huggingface.co/cuber12/dropout-decay/tree/v1.1-preprint
Publication details
- DOI
- 10.5281/zenodo.20612808
- OpenAlex
- W7164045220
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.