Poisoned Tools: Prompt-Injection Robustness and Failure Anatomy of Small Open-Weight Tool Agents
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
A reliability-first measurement of indirect prompt injection on small open-weight tool agents (1–7B, quantization fixed at 8-bit). Across 4,224 seeded episodes — 5 models × 6 training-free defenses plus a 60-template attack-surface screen, graded programmatically (blind-validated κ = 1.00 / 0.95) — four findings emerge, three against pre-registered predictions: (1) placement dominates and reverses the trusted-surface intuition, with error-string injections obeyed ~50× more than search-result bodies (0.237 vs 0.005 given exposure); (2) susceptibility rises monotonically with model size (0.020 → 0.179 → 0.236); (3) no cheap defense significantly reduces attack success, and a per-turn goal reminder costs 19 points of benign utility; (4) resistance decays under repetition (resist⁸ = 0.50 in the most-attacked cell). Harness, logs, and analysis are public and reproduce in under 4 GPU-hours at zero cost.
Publication details
- DOI
- 10.5281/zenodo.21349714
- OpenAlex
- W7168298256
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.