WhyLab: A Causal Audit Framework for Stable Agent Self-Improvement
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
Self-improving AI agents lack runtime safeguards that prevent evaluation drift, fragile outcome acceptance, and unbounded parameter updates from compounding into catastrophic policy degradation. WhyLab introduces a causal audit framework comprising three complementary defenses: C1: Information-theoretic drift detection across evaluation streams C2: E-value × Robustness Value dual-threshold filter for fragile outcomes C3: Lyapunov-bounded adaptive damping with observable energy proxy Experiments on synthetic environments demonstrate that C1 improves within-horizon detection reliability, C2 substantially reduces fragile acceptance rates, and C3 achieves the lowest violation frequency with strong proxy–state alignment. Code: https://github.com/neogenesislab/WhyLab-NeurIPS2026
Publication details
- DOI
- 10.5281/zenodo.18948929
- OpenAlex
- W7134937481
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.