CostOpt: Quantifying Cost-Accuracy Tradeoffs in LLM Prompting Strategies: An API-Only Cross-Provider Factorial Study on CostBench 500
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Current LLM reasoning strategy evaluations emphasize accuracy while ignoring cost, despite cost varying by 10$\times$ or more across configurations. We present \textsc{CostOpt}, a framework for cost-aware strategy evaluation, and \textsc{CostBench 500}, a curated benchmark of 500 static procedural tasks constructed via architecture-mixed pre-screening to eliminate zero-shot selection bias. We conduct a factorial experiment: 5 prompting strategies $\times$ 4 models at temp${}=1$ $\times$ 500 tasks $\times$ 3 seeds = 30,000 primary runs (all models at temp${}=1$) plus 15,000 supplementary ablations (Claude at temp${}=0$, Anthropic-native tool schemas). Total cost: \$441, overall accuracy: 85.9\%. All numbers from a single canonical evaluation pipeline. To ensure fair evaluation, we explicitly decouple static procedural tasks from interactive, multi-step environments. Key findings: (1) strategy choice is the dominant cost lever (Plan-and-Execute: 6.0 calls, \$0.029/run, 32.2s vs.\ Zero-Shot: 1.0 call, \$0.002/run, 4.1s); (2) on static tasks, single-shot Structured Reasoning Format achieves 89.6\% at \$0.007/run---1.3 pp higher accuracy than Plan-and-Execute at 3.9$\times$ lower cost, but this ordering completely inverts on multi-step interactive tasks where Plan-and-Execute is Pareto-optimal; (3) Structured Reasoning Format + Claude 4.5 Opus achieves 95.9\%---the best single static configuration---with Claude temp${}=0$ near-identical (mean $\Delta < 0.8$ pp); (4) GPT-5 leads tool-calling (90.9\%), GPT-5-mini leads zero-shot (88.5\%); (5) Claude tool-calling with native schemas trails GPT by 5--10 pp; (6) temperature has negligible effect on Claude accuracy. We emphasize that structural token multipliers (e.g., the 15.9$\times$ Plan-and-Execute overhead) are more durable metrics than absolute dollar costs, which remain transient. This API-only study explicitly excludes open-weights models, as their fixed CapEx and utilization-dependent OpEx require fundamentally different Pareto frontier calculations.
Publication details
- DOI
- 10.5281/zenodo.20824164
- OpenAlex
- W7165781357
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.