preprint Open access

The Enforcement Gap: Specification-Present Rationalization in Large Language Models

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Öz

Large language models demonstrate a previously undocumented failure mode in which safety and reasoning specifications are read, correctly identified, and violated in the same processing step. The model does not lack the relevant knowledge. It does not misunderstand the constraint. It identifies the rule, names the prohibition, and proceeds to violate it through motivated reasoning — reasoning paths that formally satisfy a constraint while circumventing its intended effect — driven by helpfulness optimization. This paper documents the phenomenon through reproducible testing across four major commercial systems, presents a taxonomy of three distinct failure types, and establishes through iterative specification-level testing that the failure is robust to a range of prompt-layer and specification-level remediation attempts. The solution space is characterized as architectural, operating at a level below the specification layer.

Record transparency

Publication details

DOI
10.5281/zenodo.19201964
OpenAlex
W7140226419
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.