preprint Open access

From Recursive Scaffolding to Admissibility-First Construction: Mechanism, Stability, and Failure-Mode Decomposition on OOLONG-Pairs

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
4
References
4
Comments
0
Paper overview

Öz

Empirical companion to the RLM paper, testing its theoretical predictions on a Zhang-style OOLONG-Pairs setting. Decomposes recursive scaffolding into its mechanism components — filter-mode GAF, construction-mode GAF, and the admissibility-relevant summary state — and shows empirically that the essential mechanism is not recursion itself but the construction or approximation of an admissibility-relevant state. Abstract Recursive Language Models (RLMs) have recently shown strong empirical gains on long-context reasoning and aggregation benchmarks, especially OOLONG and OOLONG-Pairs. In prior theoretical work, RLMs were analyzed through the admissibility-dynamics framework as a composite mitigation: they extend effective certification depth by offloading context into an external REPL environment, and they partially escape bounded local generation by constructing intermediate summaries through recursive sub-calls. That analysis predicts that RLMs succeed not because recursion is intrinsically sufficient, but because, in favorable cases, recursive scaffolding constructs an admissibility-relevant summary state. This paper reports an empirical decomposition of that prediction on a Zhang-style OOLONG-Pairs setting constructed from the validated trec_coarse OOLONG-synth split. In the canonical 20-query seed-7700 run at approximately 32K context tokens, direct GPT-5 collapsed under the global pair-construction burden, achieving micro-F1 of approximately 0.0019. RLM(GPT-5) recovered the global pair structure with micro-F1 = 0.9064 and macro-F1 = 0.8400. Model-based GAF filtering over RLM output increased precision and improved micro-F1 to 0.9107, although with a nontrivial recall reduction. Model-based admissibility-first GAF construction without RLM achieved the highest non-oracle micro-F1 in that run, 0.9212, but lower macro-F1 than RLM, 0.8196. This result should therefore be read not as categorical dominance over RLM, but as a different precision–recall operating point: higher recall, lower precision, and comparable aggregate F1 without recursive scaffolding. A representative five-query robustness subset (q8, q9, q11, q15, q20) was then evaluated across seeds 7701, 7702, and 7703 to probe seed sensitivity, query-level variance, candidate-generation brittleness, and model-parity concerns. Across the three-seed subset, Z1G-S-mini achieved the best practical micro-F1, 0.9169, while RLM and Z2G-M achieved 0.7304 and 0.7335 respectively. The aggregate masks an important bimodal behavior. On seeds 7701 and 7703, RLM recovered high performance (micro-F1 = 0.9048 and 0.9111). On seed 7702, however, RLM suffered a catastrophic recall collapse (micro-F1 = 0.1379), driven primarily by q8 under-generation; model-based GAF filtering inherited that collapse because it can only reject existing candidates. Construction-mode GAF did not depend on the RLM candidate set and remained high across all three seeds: Z1G-S-mini ranged from 0.9035 to 0.9244 micro-F1. The robustness result therefore separates filter-mode GAF from construction-mode GAF: filtering is valuable when a rich candidate set exists, but construction is robust to candidate-generation failure. GPT-5 construction remains a model-parity and tail-query diagnostic rather than a complete three-seed lane: across seeds 7701 and 7702 it achieved micro-F1 = 0.9217 and the strongest macro-F1, 0.7026. A direct output-structure comparison on q15 seed 7702 provides the strongest mechanism-level evidence. RLM and Z1G-S-Big produced identical aggregate counts but not identical pair sets; instead, they shared a large role-product core and differed by a symmetric single-endpoint false-positive substitution. This indicates that the recursive scaffold and direct admissibility classifier converged to the same admissibility-summary form, with disagreement isolated to endpoint-level classification noise. The updated analysis also distinguishes endpoint-classifier noise, false-positive amplification, residual pair-level relation structure, verifier miscalibration, candidate-generation collapse, benchmark/seed validity, and cost/routing as separate bottlenecks with different mitigations. The central conclusion is mechanism-level. On OOLONG-Pairs, where pair validity is often factorizable through endpoint admissibility but can expose residual pair-level constraints, recursive scaffolding is not the essential mechanism. The essential mechanism is the construction or approximation of an admissibility-relevant state. RLM is one route to that state; GAF-style admissibility-first construction states the mechanism directly. The full experimental sequence reported here remained at small-lab scale, under USD $100 in API charges in this implementation, making the bridge experiment reproducible without institutional compute. The results remain limited by adapted benchmark construction, partial multi-seed coverage, synthetic grouping, single-run conditions per seed/query, and the favorable factorizable and partially factorizable structure of OOLONG-Pairs. They should be interpreted as directionally strong empirical bridge evidence rather than a statistically powered benchmark claim. Companion Lean 4 formalization: https://doi.org/10.5281/zenodo.20062396 GitHub repository: https://github.com/shawnjason/OOLONG-Pairs Related papers in the program: PIT (foundational projection-theoretic result): https://doi.org/10.5281/zenodo.19633241NEO (forward-case impossibility theorem): https://doi.org/10.5281/zenodo.19688367IA (admissibility-dynamics framework): https://doi.org/10.5281/zenodo.19688628HAL (language-model specialization): https://doi.org/10.5281/zenodo.19715059RLM (theoretical predecessor whose predictions this paper tests empirically): https://doi.org/10.5281/zenodo.19753549SUD (Sudoku-Microscope empirical validation): https://doi.org/10.5281/zenodo.20277939HAM (Hamiltonian-Microscope cross-provider pilot): https://doi.org/10.5281/zenodo.20278073

Record transparency

Publication details

DOI
10.5281/zenodo.20277804
OpenAlex
W7161608211
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.