Decision Anchors, Fail-Closed Profiles, and Capability Tokens
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Hallucination-Aware Audit Gate (HAAG) is an observable-only protocol for audit-ready action gating in AI agents.It mediates protected actions (tool execution, publication, spending, escalation, actuation) using verifiable logs and cryptographic checks, without assuming semantic truth or model introspection.The design is post-hoc attachable: existing agents and safety modules can be integrated as measurement or policy components. HAAG enforces a strict separation between evidence events (OBS / MEASURE-COMMIT / CLAIM / POLICY-SNAPSHOT / WATERMARK) and action events (ACT). Decisions are anchored to a Decision Anchor Checkpoint in a transparency log (Merkle-tree commitments), which defines an immutable evidence prefix available at decision time. This prevents hindsight and selective evidence inclusion via anchor selection monotonicity and anchor-constrained policy selection. To reduce gaming and make attempts auditable, HAAG uses two-stage admissions (ADMIT-PRE and ADMIT-BIND) and computes anchored fingerprints (StrictFP, NormFP) scoped to the active policy snapshot at the chosen anchor. Protected actions are executed only with capability tokens that bind action parameters via an Action Digest and require execution-time verification: token signature over canonical bytes, prerequisite inclusion proofs under the anchor, expiry and replay-control posture, and optional proof-of-possession/channel binding. HAAG is intended as an enforcement substrate for practical reliability: any hallucination detection, retrieval, evaluator, or rule-based guard can be plugged in as an observable measurement module, while the gate remains implementable and auditable using standard log and cryptographic primitives.
Publication details
- OpenAlex
- W7128653215
- Document type
- preprint
- Language
- EN
- Source
- Open MIND
- Last metadata update
Comments
Log in to join the discussion.