preprint وصول مفتوح

Ceiling Effects and Convergence: Null Results for Instruction Repetition in LLM-Agent Pipelines

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

الاستشهادات
0
المراجع
0
Comments
0
Paper overview

Abstract

Context: Prompt repetition, the verbatim duplication of an input transforming <QUERY> into <QUERY><QUERY>, has been shown to improve accuracy for non-reasoning large language models on retrieval and multiple-choice benchmarks [Leviathan et al., 2025]. Objective: We ask whether applying this analogous pattern to fixed delegate instructions produces similar gains in multi-step agentic pipelines. Method: Three pre-registered controlled experiments used Claude Haiku 4.5 delegates (n=5 per condition, temperature 0.5) assigned either a single-copy or a repeated-prompt instruction under blinded binary rubric scoring, totalling 30 sessions and 3,196 messages. Results: Experiment 1 (session-ID refactoring, 6 criteria) yielded a non-significant score delta of +0.30 with five of six criteria saturated at 100% in both groups (Fisher’s p=1.000). Experiment 2 (tree-sitter scanner evaluation, 7 criteria) produced a complete ceiling effect: all 10 runs scored 7/7 (Mann-Whitney U =12.5, p=1.000). Experiment 3 (Kotlin grammar synthesis, 7 criteria) revealed rubric-runner co-design failure: three of seven criteria scored 0/1 across both groups because the required investigations were absent from the runner prompt. On four reachable criteria, control scored a mean of 2.00/4 and treatment scored 2.40/4 (U =15, p=0.607). Across all three pilot experiments, we detect no effect of prompt repetition on task success. Treatment agents used 30.6% fewer total tokens in Exp1 but 7.2% and 19.9% more in Exp2–3; the direction reverses with session length, as long sessions save turns while short sessions pay pure overhead. Conclusion: Criteria requiring investigations absent from the runner prompt are unreachable by both groups and must be resolved before effect sizes are interpretable This is a preprint; it has not undergone peer review.

Record transparency

Publication details

DOI
10.5281/zenodo.21795787
OpenAlex
W7172470284
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.