preprint Open access

Defense-in-Depth AI Governance: Combining Pre-Inference and Post-Inference Detection for Robust Safety Assurance

  • Open MIND
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Öz

We present a defense-in-depth architecture combining two complementary proprietary governance systems — PatternWall (pre-inference adversarial prompt detection) and Sensus (post-inference multi-dimensional output evaluation) — operating at distinct points in the inference pipeline. We benchmark the combined system across three frontier models (Claude Opus 4.6, GPT-5.2, Grok 4.1) using a custom 12-sequence red team corpus and a curated 120-task subset of the CyberGym exploit generation benchmark. Model-native safety ranges from 0.0% to 81.7% depending on model and attack type. The combined governance system achieves 77.1–91.4% detection on multi-turn social engineering and 76.7–98.3% on single-turn exploit generation. A bidirectional feedback loop between layers recovers up to 5 additional adversarial turns on multi-turn attacks, improving detection from 77.1% to 91.4%. These findings establish that external, model-agnostic governance middleware provides consistent safety assurance regardless of underlying model behavior. Patent pending (LKM-2026-001). Preprint submitted to SSRN.

Record transparency

Publication details

DOI
10.5281/zenodo.18940959
OpenAlex
W7134828582
Document type
preprint
Language
EN
Source
Open MIND
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.