article Open access

Adversarial Baseline Poisoning: A Novel Attack Class Against Behavioral Scoring Systems for AI Agents

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

AI agents deployed in enterprise environments are increasingly governed by behavioral scoring systems that detect anomalous activity by comparing current behavior against historical baselines. We identify and formally characterize a novel attack class — adversarial baseline poisoning (ABP) — in which a compromised agent gradually shifts its own behavioral baseline over time, causing a behavioral scoring system to accept progressively malicious behavior as normal. We demonstrate that score-only detection fails against this attack class (20% detection rate on a synthetic financial services corpus), and present a two-layer defense architecture that achieves 100% detection across ten multi-day attack scenarios at 0.00% false positive rate on a 150-scenario legitimate behavior corpus (bounding FP below 2.4% at 95% CI, Clopper-Pearson). The defense is implemented in the open-source AgentRepEngine runtime enforcement system.

Record transparency

Publication details

DOI
10.5281/zenodo.19535881
OpenAlex
W7153891693
Document type
article
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.