Logo image
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
Preprint   Open access

The Price of Thinking: Reasoning Effort as a Model-Specific API Contract

Yeabin Moon
arXiv (Cornell University)
08/16/2026

Abstract

Large language models Reasoning effort AI model evaluation AI API pricing Cost–accuracy tradeoffs Reproducible benchmarking
API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls. Mean delivered cost was $0.01031 per call higher under the explicit-high contract than under the omitted contract [+$0.00204, +$0.01974]. The corresponding accuracy contrast was +0.0133 [-0.0267, +0.0467]; we did not detect an accuracy difference, and the interval permits a gain of up to 4.67 percentage points that this design cannot rule out. Cost per correct answer was $0.08665 under the high-effort contract and $0.07662 under the omitted contract, as registered point estimates. A dated contract census, Models-API metadata, and preregistered raw-response probes further documented model-specific omission semantics, including within a provider; claims remained at documentation grade when raw structure was indeterminate. The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined; the resulting claims are bounded to the model, task, and collection date studied.
pdf
2608.16956v1403.25 kBDownloadView
Open Access
url
https://arxiv.org/abs/2608.16956View

Metrics

1 Record Views

Details

Logo image