article Open access

Binary Bitmask Verification vs LLM Self-Review at Service Scale: The Robustness Verification Framework

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
1
Comments
0
Paper overview

Öz

At service scale ≥ 200 components, interpretation-based verification fails. LLM self-review depends on the writer-agent reviewing its own output — structurally insufficient when the agent cannot perceive the defect it introduced. We present the Robustness Verification Framework (RVF), a four-layer binary-first verification stack that subordinates LLM judgment to deterministic bitmask predicates. Empirical: deployed across 261 services in the ANKR substrate over four sessions in May 2026; surfaced 660 silent bugs across 175 source files (the "botched latency bulk-edit" pattern), 11 codex overclaims, 22 R-008 port-authority violations, 2 invalid capability_captured claims, and one fleet-wide gitleaks-quirk; 19 services formally proof-eligible per the 6-bit Layer C threshold; 5/5 negative-gate scenarios pass; 0 false-positive production blocks. We argue capability hallucination at scale is structural to LLM-author/LLM-reviewer systems and that binary verification is the only enforceable response. Companion to ArchMask (10.5281/zenodo.19417112), BitMaskOS (10.5281/zenodo.20047909), PRAMANA Protocol (10.5281/zenodo.19273330), and Vivechana SAR-006 (10.5281/zenodo.19456053).

Record transparency

Publication details

DOI
10.5281/zenodo.20200378
OpenAlex
W7161282248
Document type
article
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.