preprint Open access

CostOpt: Quantifying Cost-Accuracy Tradeoffs in LLM Prompting Strategies: An API-Only Cross-Provider Factorial Study on CostBench 500

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Current LLM reasoning strategy evaluations emphasize accuracy while ignoring cost, despite cost varying by 10$\times$ or more across configurations. We present \textsc{CostOpt}, a framework for cost-aware strategy evaluation, and \textsc{CostBench 500}, a curated benchmark of 500 static procedural tasks constructed via architecture-mixed pre-screening to eliminate zero-shot selection bias. We conduct a factorial experiment: 5 prompting strategies $\times$ 4 models at temp${}=1$ $\times$ 500 tasks $\times$ 3 seeds = 30,000 primary runs (all models at temp${}=1$) plus 15,000 supplementary ablations (Claude at temp${}=0$, Anthropic-native tool schemas). Total cost: \$441, overall accuracy: 85.9\%. All numbers from a single canonical evaluation pipeline. To ensure fair evaluation, we explicitly decouple static procedural tasks from interactive, multi-step environments. Key findings: (1) strategy choice is the dominant cost lever (Plan-and-Execute: 6.0 calls, \$0.029/run, 32.2s vs.\ Zero-Shot: 1.0 call, \$0.002/run, 4.1s); (2) on static tasks, single-shot Structured Reasoning Format achieves 89.6\% at \$0.007/run---1.3 pp higher accuracy than Plan-and-Execute at 3.9$\times$ lower cost, but this ordering completely inverts on multi-step interactive tasks where Plan-and-Execute is Pareto-optimal; (3) Structured Reasoning Format + Claude 4.5 Opus achieves 95.9\%---the best single static configuration---with Claude temp${}=0$ near-identical (mean $\Delta < 0.8$ pp); (4) GPT-5 leads tool-calling (90.9\%), GPT-5-mini leads zero-shot (88.5\%); (5) Claude tool-calling with native schemas trails GPT by 5--10 pp; (6) temperature has negligible effect on Claude accuracy. We emphasize that structural token multipliers (e.g., the 15.9$\times$ Plan-and-Execute overhead) are more durable metrics than absolute dollar costs, which remain transient. This API-only study explicitly excludes open-weights models, as their fixed CapEx and utilization-dependent OpEx require fundamentally different Pareto frontier calculations.

Record transparency

Publication details

DOI
10.5281/zenodo.20824164
OpenAlex
W7165781357
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.