The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
06:00 · August 19, 2026 · arXiv cs.AI RSS

API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls. Mean delivered cost was \$0.01031 per call higher under the explicit-high contract than under the omitted contract [+\$0.00204, +\$0.01974]. The corresponding accuracy contrast was +0.0133 [-0.0267, +0.0467]; we did not detect an accuracy difference, and the interval permits a gain of up to 4.67 percentage points that this design cannot rule out. Cost per correct answer was \$0.08665 under the high-effort contract and \$0.07662 under the omitted contract, as registered point estimates. A dated contract census, Models-API metadata, and preregistered raw-response probes further documented model-specific omission semantics, including within a provider; claims remained at documentation grade when raw structure was indeterminate. The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined; the resulting claims are bounded to the model, task, and collection date studied.
Summary
API buyers acquire a dated contract rather than a model name in isolation. That contract bundles the requested and served model, the reasoning-effort setting or its omission, output constraints, service tier, prompt, and prevailing price schedule. The study isolates the effect of one such term by running a preregistered, paired comparison of Claude Sonnet 5 under explicit high effort against the identical model with the effort parameter left unspecified. Both conditions default to the provider’s documented high-effort and adaptive-thinking behavior; the contrast therefore measures the difference between stating the parameter and relying on the default.
The evaluation used the full set of 30 AIME 2026 problems, each presented five times under each contract. Every completed call received a single terminal label drawn from a fixed taxonomy that distinguishes correct answers, incorrect answers, rail-exhausted refusals, other non-answers, and provider failures. Item-level resampling preserved the repeated calls while treating problems as the clustering unit. Mean delivered cost rose by $0.01031 per call when high effort was stated explicitly, with the registered interval running from $0.00204 to $0.01974. The corresponding accuracy difference of +0.0133 fell short of statistical significance; the interval allows for a possible gain as large as 4.67 percentage points but does not confirm one.
Point estimates of cost per correct answer were $0.08665 under the explicit-high contract and $0.07662 under the omitted contract. Supporting analyses of model metadata and raw responses further showed that omission semantics vary across models, even within a single provider, so that provider-level descriptions alone are insufficient for cost forecasting. The design froze the request registry, parser, outcome taxonomy, and analysis pipeline before any outcome data were examined, limiting claims to the specific model, task, and collection window studied.
Why it matters
This research is highly relevant for Dutch AI researchers and MLOps practitioners focused on cost-efficient AI deployment. Understanding the hidden costs and stochastic nature of reasoning API contracts enables Dutch SMEs and enterprises to optimize their AI infrastructure and routing strategies.










