Reasoning effort entered LLM products disguised as an ordinary control: low, medium, high. It looks like a quality setting. It is closer to a budget. Raising the value does not make the model more knowledgeable; it changes the policy governing how much inference-time work the same model can spend before it answers.
On a single request, that distinction may seem academic. Across a long-running agent, it governs hidden generation, latency, cache reuse, and how much of a finite budget is spent planning rather than acting. Raising effort indiscriminately wastes compute. Holding it too low can starve the few decisions on which the rest of a run depends.
A few years ago, developers tried to elicit reasoning by writing “think step by step,” supplying worked examples, or sampling several answers and voting. Then models were trained to produce useful intermediate work. Providers began charging for hidden reasoning tokens. Open-weight releases exposed the training recipes. Agent harnesses learned to vary the budget from one step to the next. What began as a prompting technique has become a layer of inference infrastructure, often still treated as a UI preference.
Reasoning effort is not an intelligence slider. It is an inference-time compute policy — and the right default is the lowest level that makes the system work reliably.
To understand what the selector means—and why providers and agent harnesses cannot simply hide it forever—we need to follow the path by which reasoning became something engineers could budget.
How reasoning became a budget #
2022: reasoning was something you prompted for
The modern story starts before there was an effort field. The chain-of-thought paper showed that sufficiently large language models could solve some multi-step problems more accurately when prompts included worked intermediate reasoning. Self-consistency then showed a second route to better results: sample several different reasoning paths and select the answer they most often reach. One method spent more tokens inside a trajectory; the other spent more compute across trajectories. Both established the idea that inference could be scaled after training.2324
At this stage, the developer owned the machinery. A prompt encouraged a trace, a sampling loop created alternatives, and application code chose the result. “Reasoning effort” was an emergent consequence of prompting and decoding, not a calibrated model capability.
2023: the process became a training target
The next step moved reasoning from prompt craft into post-training. Work on process supervision compared rewarding correct intermediate steps with rewarding only the final outcome. In mathematics, OpenAI reported that process-supervised reward models selected correct solutions more reliably than outcome-supervised ones as more candidate solutions were considered. This did not settle the general training recipe, but it made a central design question explicit: should a model learn from the path, the destination, or both?25
Verifiable domains such as mathematics and code offered another possibility. If a checker can determine whether the final answer or program is correct, reinforcement learning can reward successful behavior without requiring a human to annotate every hidden step. That idea would become central to the open reasoning-model wave.
September 2024: test-time compute became a product
OpenAI’s o1 release was the moment the shift became visible to ordinary API users. OpenAI described a model trained with reinforcement learning to refine its chain of thought, recognize mistakes, and try other approaches, and reported that performance improved with both train-time compute and time spent thinking at inference. Raw chain-of-thought was withheld, while a model-generated summary could be shown. The internal trajectory had become a metered product behavior rather than prompt text the application necessarily owned.26
This changed the engineering question. Developers no longer asked only, “How should I prompt the model to reason?” They also had to ask, “How much invisible generation should this request be allowed to consume, how do I measure it, and when does the extra latency pay off?”
January 2025: open weights exposed the recipe
DeepSeek-R1 made the training story inspectable. R1-Zero applied large-scale reinforcement learning before supervised fine-tuning and developed longer reasoning, reflection, and self-correction. The full R1 pipeline added cold-start data, further training for general usefulness, and distillation into smaller models. The release turned techniques that had largely been inferred from closed systems into code, weights, and a detailed report.8
2025: reasoning became configurable
Once reasoning was a model behavior, providers began turning it into a family of controls. Qwen3 offered hybrid thinking and non-thinking modes. OpenAI’s gpt-oss encoded low, medium, and high effort in the Harmony prompt protocol. Hosted APIs exposed qualitative levels, numeric budgets, or adaptive modes. Gateways such as OpenRouter normalized the wire format, while hosts such as Baseten surfaced the native semantics of individual open models.91067
2026: one label, several mechanisms
The current generation has made the term effort less uniform, not more. It can mean a learned ordinal mode, a hard token budget, a continuous conditioning value, an adaptive provider policy, preserved thinking across tool calls, or—in one multi-agent API—the number of collaborating agents. Agent harnesses now sit above those mechanisms, allocating compute across planning, acting, verification, retries, and handoffs.
September 2026: effort becomes mutable conversation state
The next step arrived almost quietly. Anthropic documented per-message effort for Claude on 1 September; OpenAI followed with GPT-6 Astra on 3 September. Both let an application append a privileged configuration event that changes effort for later turns without rewriting the conversation that came before it. A control that used to belong near the beginning of a request had become mutable state inside a long-running conversation.1431
That history explains the present confusion. The industry converged on the need to control inference-time work, but not on a unit for measuring it. The rest of this article follows the stack downward—from the semantics of the control, through tokens and model training, and back upward into agent policy and production evaluation.
What reasoning effort actually controls #
A conventional description of an LLM request has two parts: tokens go in and tokens come out. A reasoning model adds a meaningful middle stage. Before and sometimes between pieces of visible output, the model can spend computation deciding how to approach the task. It may decompose the problem, keep track of intermediate results, reconsider an assumption, decide which tool to call, or inspect whether a proposed answer is internally consistent.
The effort setting influences that middle stage. It is generally a policy signal, not a reservation for an exact number of tokens or seconds. A capable model can use less reasoning on an easy request even when a high setting is available, and more on a difficult request at the same setting. OpenAI describes its effort control as guidance for how many reasoning tokens a model should generate, while Anthropic and Google now document explicitly adaptive thinking modes.135
This explains why the effect is uneven. Raising effort for a deterministic format conversion may produce no observable benefit because there is no useful search to perform. Raising it for a subtle concurrency bug may let the model trace an execution order it would otherwise miss. The control is valuable when the task has branches, dependencies, ambiguity, or a need for verification.
More opportunity is not always more value
Additional reasoning has diminishing returns. At some point the model has already found the relevant approach, and more work only repeats or elaborates it. Effort also cannot substitute for missing information: a model without the right context or tool will not reason its way out of that gap.
Higher effort can also make a response worse. The model may overcomplicate a direct question, invent immaterial edge cases, or spend its output budget on an unproductive line of reasoning. Treat effort as a resource allocation decision, not a one-way quality upgrade.
Reasoning tokens are part of the budget #
Providers expose internal reasoning differently, but the work has operational consequences even when it is not readable. In OpenAI’s Responses API, reasoning tokens are not returned as raw chain-of-thought. They still occupy space in the context window, appear in usage accounting, and are billed as output tokens. Their count is available under output_tokens_details.reasoning_tokens.1
| Token class | Visible to the user? | Consumes request budget? |
|---|---|---|
| Input tokens | Usually | Yes |
| Reasoning tokens | No raw trace | Yes |
| Visible output tokens | Yes | Yes |
The generated-token limit therefore covers more than the prose a user sees. If a model consumes the available envelope while reasoning, a request can finish with an incomplete status before producing a useful visible answer. This failure is surprising when an application assumes that max_output_tokens is simply a maximum answer length.
When first testing a reasoning model, leave enough space for both internal work and the final response, then inspect actual usage. The stable engineering practice is to measure the distribution for representative tasks and set limits from observed behavior. A single average hides exactly the requests most likely to time out.
Reasoning effort is not verbosity #
Several API controls influence a response, and their effects are easy to collapse into one vague idea of “more.” Keeping them separate makes experiments much easier to interpret. Reasoning effort changes the internal work used to solve a task. Verbosity changes how much of the answer is explained. The output-token limit defines a hard ceiling. Temperature changes sampling variability. The context window limits how much information can participate in a request.
| Control | Primarily changes | Does not guarantee |
|---|---|---|
reasoning.effort | Internal problem-solving work | A longer or correct answer |
| Verbosity | Detail in the visible response | Deeper analysis |
| Output-token limit | The hard generation ceiling | Efficient use of the ceiling |
| Temperature | Sampling variability | More careful reasoning |
| Context window | Information available to the model | Attention to every detail |
A model can think extensively and return three sentences. It can also produce two pages of confident prose after shallow reasoning. For this reason, visible response length is a poor proxy for computational work. Asking a model to “explain in detail” may improve communication, but it is not equivalent to increasing reasoning effort.
The same separation is useful in product design. If users want a concise answer to a difficult question, configure more reasoning and lower verbosity. If they want a tutorial for a simple concept, moderate reasoning and higher verbosity may be the better combination. Treat the controls as independent until an evaluation shows an interaction that matters for your workload.
The control is not standardized #
“Reasoning effort” is a product-level name for several different controls. Some APIs expose qualitative levels. Some expose a numeric thinking-token budget. Some let the model decide dynamically. A gateway may translate one vocabulary into another, but that translation does not make the underlying compute equivalent. Code should validate capabilities per model and log the provider-native setting that was actually sent.
| Surface | Primary control | Important behavior |
|---|---|---|
| OpenAI Responses | reasoning.effort | Qualitative levels; reasoning tokens share the generated-output envelope and raw chain-of-thought is not returned. |
| Anthropic Messages | thinking.type plus output_config.effort | Adaptive thinking can decide whether and how deeply to think; effort also affects text and tool-call token use. |
| Google Gemini | thinking_level | Dynamic thinking is the default on current thinking models; supported levels and whether thinking can be disabled vary by model. |
| xAI Responses | reasoning.effort | On Grok 4.6 it changes reasoning depth; on the Grok multi-agent model it changes how many agents collaborate. |
| OpenRouter | reasoning.effort or reasoning.max_tokens | A normalized gateway surface translated to provider-native levels or budgets. |
| Baseten Model APIs | Model-specific chat-template or OpenAI-compatible fields | Open-weight models keep their native semantics; for DeepSeek V3.2, thinking is enabled with a chat-template argument. |
Reference appendix: provider notes and API examples
OpenAI: qualitative effort
OpenAI’s Responses API places the control inside reasoning. Supported values and defaults depend on the model. The usage object separates reasoning tokens from visible output tokens, which makes an effort sweep observable even though the raw reasoning trace remains hidden.1
const response = await client.responses.create({
model: "gpt-5.6",
reasoning: { effort: "medium" },
input: "Review this migration plan for data-loss risks.",
max_output_tokens: 32000
});
console.log(response.output_text);
console.log(response.usage.output_tokens_details.reasoning_tokens);
Anthropic: thinking mode and response effort
Anthropic separates the decision to think from the effort applied to the whole response. On models that support adaptive thinking, thinking: {type: "adaptive"} lets Claude decide whether and how much to think, while output_config.effort steers the overall token expenditure. Older extended-thinking models instead use a fixed budget_tokens; the two forms are not interchangeable across model generations.34
response = client.messages.create(
model="claude-opus-4-8",
max_tokens=16000,
thinking={"type": "adaptive"},
output_config={"effort": "medium"},
messages=[{"role": "user", "content": prompt}],
)
Gateways and open-model providers
xAI’s current Responses API illustrates why field names alone are insufficient. For Grok 4.6, reasoning.effort selects low, medium, high, or xhigh reasoning depth, and reasoning cannot be disabled. On grok-4.20-multi-agent, the same field controls how many agents collaborate instead of the depth of one model trajectory. Usage policy and orchestration policy can therefore share a wire shape while spending compute in fundamentally different ways.21
OpenRouter offers a normalized reasoning object. A caller can request an effort label or, where supported, a maximum reasoning-token budget. The gateway maps that request to the selected provider’s native control and publishes per-model capabilities through its models endpoint. This is useful portability, but it is translation rather than standardization: an OpenRouter medium request can become a provider level on one model and a budget derived from the output limit on another.6
{
"model": "your-model",
"messages": [{"role": "user", "content": "..."}],
"reasoning": {"effort": "high"}
}
Baseten exposes managed open-weight models through an OpenAI-compatible endpoint, but reasoning controls remain model-specific. Its DeepSeek V3.2 example enables thinking with chat_template_args.enable_thinking and returns the trace separately as reasoning content. That is a switch, not a portable low-to-high ladder. A self-hosted deployment can expose still different knobs depending on its tokenizer, chat template, and serving engine.7
response = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V3.2",
messages=[{"role": "user", "content": prompt}],
extra_body={"chat_template_args": {"enable_thinking": True}},
)
Across all of these APIs, record four things: the application’s intended effort, the exact provider-native request, observed usage, and the model version. Without that separation, a gateway migration can silently change the meaning of an experiment.
Inside open-weight reasoning models #
Open-weight models make the mechanism easier to inspect. In the common design, “thinking” is not a second symbolic engine bolted onto a transformer. The model autoregressively generates an intermediate token sequence—often in a dedicated channel or between special tags—and then generates the answer conditioned on that sequence. More reasoning effort usually means allowing or encouraging a longer intermediate trajectory, not giving each transformer token a fundamentally different computation.
There are two separate engineering problems. Training must teach the model that useful intermediate work improves the eventual answer. Serving must format the request correctly, let the model emit and reuse that work, decide when to stop it, and keep the reasoning channel separate from user-visible text. A model can have strong learned reasoning behavior while a serving stack exposes only an on/off switch; conversely, a server can expose a reasoning_effort field that the checkpoint was never trained to interpret.
Training scaling and inference scaling are different axes
Training a larger or better post-trained model changes the weights. Increasing reasoning effort keeps the weights fixed and spends more computation during generation. A smaller model at high effort can sometimes overlap a larger model at low effort, but neither dominates universally: model capacity and inference budget are separate variables that must be swept together.
A longer single reasoning trace is only one form of inference scaling. A harness can sample several independent solutions and select by majority vote or a verifier, ask the model to critique and revise an answer, search a tree of candidate steps, or use several agents and synthesize their results. These methods spend compute across trajectories rather than only extending one trajectory. They can outperform “turn effort to max,” but they also add orchestration cost and need a trustworthy selection signal.
How models learn controllable effort
DeepSeek-R1 popularized reinforcement learning with verifiable rewards: for domains such as math or code, a checker can reward a correct final answer without requiring a human-written reasoning trace. The model learns longer scratch work, backtracking, and self-correction because those behaviors increase the probability of receiving the outcome reward. The literal <think> tags are just delimiters — the capability comes from training; the delimiter makes the protocol operable.8
Current open-weight reports disclose four recurring ways to make effort controllable:
- Mixed supervised fine-tuning. Train the same checkpoint on short direct answers and longer reasoning examples, conditioned on a mode or effort instruction.
- Mode-conditioned reinforcement learning. Change the per-token cost, length penalty, context allowance, or reward constraint according to the requested effort.
- Specialists followed by distillation. Train separate domain or effort specialists and distill their behaviors into one checkpoint that can switch modes at inference.
- Budget-aware continuation. Train or verify that the model can produce a useful answer after its reasoning trace is truncated or externally closed.
A plain-language system instruction works only when post-training has taught the model to interpret it. Adding “reason at maximum effort” to an arbitrary instruct model may change its prose, but it does not manufacture a calibrated effort mode.
DeepSeek-R1: what inference looks like
At inference time, R1's <think>...</think> structure is simply generated text with special semantics in the chat template and parser. The reasoning trace consumes output tokens before the final answer. A host can cap total generation, stop or hide the reasoning segment, or retain it for the next turn. Those choices affect cost and behavior even though none changes the trained weights.8
Qwen3: one checkpoint, two modes
Qwen3 was trained for hybrid thinking and non-thinking behavior. Its official Transformers instructions enable thinking by default and allow a hard switch through enable_thinking=False; when thinking is enabled, /think and /no_think can steer individual turns. The chat template renders the appropriate prefix and the generated reasoning is parsed away from the final answer.9
This is an instructive contrast with a numerical token budget. A mode switch changes the behavior the model is prompted and trained to follow. It does not promise that every thinking request consumes the same number of tokens, nor that a non-thinking response performs zero internal computation. Qwen3 also supports a hard reasoning budget: the runtime can stop the trace near a threshold, insert a stop-thinking instruction, and continue to the answer. Its technical report says useful continuation under that intervention emerged after Thinking Mode Fusion rather than from explicit truncation training.
/think and /no_think examples, empty thinking tags included. From Raschka’s survey.gpt-oss: effort in the prompt protocol
OpenAI’s gpt-oss models are mixture-of-experts open-weight reasoning models with configurable low, medium, and high effort and an accessible chain-of-thought. They were trained on the Harmony response format, which places the effort instruction in the system message and separates analysis, tool-call commentary, and final-answer channels. A runtime such as vLLM or a hosted provider may render and parse Harmony for you; a custom inference loop must do so correctly.10
Reasoning: high inside the opening Harmony system message. From Raschka’s survey.Reference appendix: what the 2026 open-weight releases added
The newest open-weight families make the design space more explicit. The names are still inconsistent, but their reports now describe how learned modes, token costs, specialists, hard budgets, and multi-turn reasoning state fit together.
| Model | Inference control | Training or serving idea |
|---|---|---|
| DeepSeek V4 | Non-think plus low, high, and max thinking effort | Mode-specific post-training; higher modes use different prompting and budget policy before behaviors are combined. |
| Nemotron 3 Ultra | Reasoning mode plus an external token budget | Learned modes are paired with budget-aware training so the server can close a reasoning span near a limit and continue to the answer. |
| Kimi K2.5 | Thinking or instant mode | Token-Efficient RL alternates budgeted and unconstrained phases so shorter reasoning does not destroy test-time scalability. |
| Kimi K3 | low, high, max | Effort specialists across general, coding, and agentic domains are combined into one model with multi-teacher on-policy distillation. |
| Qwen3.8 | Explicit reasoning effort plus thinking on/off | Adds preserved thinking so prior reasoning state can participate in long tool-using conversations. |
| GLM-5.3 | low, high, max; thinking always on | Builds on interleaved, preserved, and per-turn thinking for long tool-using runs. |
| Inkling | Continuous effort rather than a few labels | The effort value is placed in the system message while the per-token cost changes during large-scale RL. |
DeepSeek’s current V4 API accepts low, high, and max, while non-thinking is selected separately. Its released encoding shows that the reasoning level changes a text prefix before the system message. This is a useful warning for gateways: translating a generic medium to a model without a native medium mode requires an explicit mapping, and third-party hosts must implement the checkpoint’s template faithfully.13
Nemotron 3 Ultra separates a learned reasoning mode from an inference-time budget. The model’s public materials describe reasoning-budget control, SFT and reinforcement learning, and multi-teacher on-policy distillation. In the budgeted serving pattern, the client can request that reasoning end near a threshold; budget-aware training helps the model transition from a partial trace to a final answer rather than collapse when the trace is cut short.14
Kimi K2.5’s Token-Efficient RL addresses a different failure: training only against short fixed budgets can make a model concise but unable to benefit from extra compute later. Its Toggle method alternates budgeted phases, where correct solutions are encouraged to stay under a problem-specific limit, with unconstrained phases that restore the normal generation ceiling. The released policy has no effort selector for these phases—the method makes its default thinking policy more token-efficient while retaining test-time scaling.20
Kimi K3 exposes three effort levels and requires preserved reasoning content in multi-turn and tool-calling conversations. Its technical report trains specialists for general, coding, and agentic work at multiple effort budgets, then combines them into one model. That is a stronger form of effort conditioning than merely prompting one generic policy to be shorter.15
Qwen3.8 extends Qwen’s hybrid modes with explicit reasoning-effort control and preserved thinking history. The latter is especially relevant to agents: retaining prior reasoning can improve continuity and cache reuse, but it also grows context and requires the harness and inference runtime to agree on the exact chat-template contract.22
GLM’s progression emphasizes agent state. Interleaved thinking lets the model reason between tool calls, preserved thinking retains prior reasoning blocks, and turn-level control changes thinking per request. The August 2026 GLM-5.3 release adds low, high, and max effort and removes the thinking-off mode. A harness migrating from older GLM versions must therefore map “off” to low effort rather than sending a now-invalid disabled setting.16
Thinking Machines Lab’s Inkling replaces ordinal labels with a continuous effort value. During its large-scale asynchronous RL, the requested effort appears in the system message and changes the cost assigned to generated tokens: a higher token cost encourages brevity, while a lower cost permits longer reasoning. This demonstrates a general principle behind effort-conditioned RL without implying that every provider uses the same recipe.17
These examples reveal the core pattern: effort is part learned policy, part prompt protocol, and part serving policy. For open weights, you can modify every layer—fine-tune the behavior, change the template, force or suppress thinking, cap generation, or build a controller that asks the model to continue. That control is powerful, but it also means the API contract is yours to test.
Why not choose effort automatically? #
Providers can, and increasingly do. Gemini documents dynamic thinking as the default on current thinking models. Anthropic’s adaptive mode lets Claude decide whether a turn needs thinking and how deeply to pursue it. OpenAI describes reasoning effort as guidance rather than a fixed token reservation. In all three cases the model can respond directly to an easy prompt and spend more tokens on a difficult one within the permitted policy.135
That still does not solve the application’s optimization problem. The provider sees the prompt and conversation, but it may not know that a wrong database migration recommendation can destroy data, that this user is waiting behind a 900 ms latency target, that a deterministic validator will catch a bad extraction, or that the next tool call moves money. “How difficult is this prompt?” is only one input. The real question is “How much should this system spend, given the value and risk of this step?”
An agent harness has more of that context, but it also works from imperfect signals. Prompt length is a poor proxy for difficulty. A classifier can misroute novel cases. A validator can approve an answer that is plausible but wrong. Fully automatic routing therefore becomes a policy learned or written around observable task features, followed by evaluation—not a free replacement for choosing an effort level.
The most useful division of responsibility is hierarchical. The application chooses an envelope from business context: model, maximum effort, token budget, deadline, and whether escalation is allowed. The provider or model adapts within that envelope to the prompt. The harness observes the result and decides whether to accept, retry, escalate, or ask for human review.
The role of the agent harness #
An agent harness owns the loop around model calls: state, tools, retries, handoffs, limits, and observability. Reasoning effort is therefore not merely a startup option on the agent. It can be a per-call decision that changes as the run moves from classification to planning, tool execution, verification, and final synthesis.
What a harness can control
- Per-step effort. Use low effort for intent classification or schema extraction, then high effort for an irreversible plan review.
- Model routing. Move a difficult step to a stronger model instead of increasing effort on a model that has reached its capability ceiling.
- Escalation. Start cheaply, run a validator, and retry at higher effort only after a concrete failure or low-confidence result.
- Budgets. Cap tokens, model calls, tool calls, elapsed time, and spend across the whole run rather than one response.
- Continuity. Preserve provider-specific reasoning items or thinking blocks across turns when the API requires them; do not reduce history to visible prose blindly.
- Observation. Attribute tokens, latency, errors, and quality outcomes to the step and policy that caused them.
Frameworks expose pieces of this control at different layers. The OpenAI Agents SDK accepts reasoning settings on an agent or run and aggregates reasoning-token usage across all model calls. LangChain middleware can intercept each model call and replace the model or its configuration from current state, which is the natural hook for task-aware effort routing.1112
function policy(step, state) {
if (state.deadlineMs < 1200) return { model: "fast", effort: "low" };
if (step.kind === "classify") return { model: "fast", effort: "low" };
if (state.validatorFailures > 0) return { model: "strong", effort: "high" };
if (step.isIrreversible) return { model: "strong", effort: "high" };
return { model: "default", effort: "medium" };
}
This policy is deliberately legible. A learned router may eventually outperform it, but explicit rules are easier to audit while the workload is young. Log the rule that fired, the provider-native request, total run budget before and after the call, and the validation outcome. That record turns “the agent thought too much” into a debuggable system event.
Reasoning state is part of agent state
Some providers return reasoning summaries, encrypted reasoning items, signed thinking blocks, or plain reasoning text. Their continuation rules differ. The harness must preserve the structures required for tool-use continuity and multi-turn reasoning while still avoiding accidental disclosure in logs or user-visible messages. Treat reasoning state as typed protocol data, not as an ordinary assistant string.
When changing effort stops breaking the cache #
Per-step effort sounds economical until the conversation becomes long. Many model protocols serialize effort near the beginning of the prompt. Changing low to high then changes an early token, so the previously cached prefix no longer matches. The harness saves reasoning tokens on one turn and pays to process the entire history again on the next.
GPT-6 Astra and recent Claude models change that trade-off. They let the application append an effort update to an existing conversation. The old history remains byte-for-byte unchanged and eligible for prompt-cache reuse; only the new control event and later turns have to be processed under the new policy. Qwen3 offered an earlier, less formal open-weight analogue, while Gemini can preserve the economics of a stable cached corpus without making the same arbitrary-conversation guarantee.
Reasoning effort is becoming mutable runtime state. The important shift is from rebuilding a prompt to appending a privileged event.
A strict definition
A model qualifies here only when all four technical conditions hold: a conversation already has a cacheable prefix; effort can change for a later turn; the earlier serialized prefix is not rewritten; and the serving API or runtime can reuse that prefix computation. A provider-documented guarantee makes the claim exact. Merely supporting prompt caching and an effort parameter is not enough.
Why one early change invalidates everything after it
A causal transformer computes each token from all the tokens before it. If effort is encoded at the front of the sequence, changing the setting changes the input to every later cached state:
# Cache-breaking
[effort=low] [system] [turn 1] [turn 2] ...
[effort=high] [system] [turn 1] [turn 2] ...
# Cache-preserving
[system] [turn 1] [turn 2]
[privileged control: effort=high] [turn 3]
If every turn adds roughly Δ tokens and the full history is prefetched again, the processed prefix grows like Δ(1 + 2 + … + n): roughly quadratic in the number of turns. Reusing unchanged prefixes moves the newly prefetched text toward O(nΔ). This reduces time to first token, input processing, and accelerator work. It does not remove the cost of decoding new reasoning or visible output, and actual savings still depend on cache retention, routing, and provider pricing.2932
The exact implementations: Astra and Claude
GPT-6 Astra adds a configuration_update item to the Responses API. The update applies to the next user turn, persists until another update, and currently changes reasoning effort only. OpenAI documents that the existing prompt prefix remains cached. The feature is limited to Astra's standard single-agent configuration; adjacent update items are rejected; and it is not directly compatible with automatic compaction, automatic truncation, or a standalone compact request. After an explicit compaction trigger, the application must reapply the desired effort.131
{
"type": "configuration_update",
"reasoning": { "effort": "high" }
}
Claude Fable 5.1, Mythos 5.1, and Opus 5 use a similar idea through an appended empty system message containing output_config.effort. With the mid-conversation-output-config-2026-07-01 beta header, the change starts with the next user turn, persists until overridden, and leaves earlier messages unchanged. Anthropic explicitly warns that changing the traditional top-level effort parameter still restarts the cache. The beta header's date is a protocol identifier, not a release date; public release notes announced per-message effort on 1 September 2026.433
{
"role": "system",
"content": [],
"output_config": { "effort": "high" }
}
Claude's direction is broader than effort: Anthropic also supports mid-conversation changes to system instructions and tool definitions. That makes this pattern especially relevant to long-running agents whose execution policy changes while their semantic history should remain stable.33
Current landscape: exact, analogous, and cache-breaking implementations
| Model | Where effort is encoded | Across-change cache status | Classification |
|---|---|---|---|
| GPT-6 Astra | Appended configuration_update | Provider-documented | Exact |
| Claude Fable 5.1, Mythos 5.1, Opus 5 | Appended system event with output_config.effort | Provider-documented beta | Exact |
| Qwen3 | Latest message can contain /think or /no_think | Possible with runtime prefix caching | Open-weight analogue |
| Gemini | Per-request thinking plus named cached content | Stable cached base is reusable; arbitrary dialogue transition is not guaranteed | Economic near-equivalent |
| Grok 4.6 | Top-level effort plus automatic prefix caching | Not documented across an effort change | Unconfirmed |
| Kimi K3 | Top-level effort selected for the conversation | No documented mid-conversation transition | Not equivalent |
| DeepSeek V4 | Effort prefix before the system message | Early tokens change | Cache-breaking |
| GLM-5.3 | Instruction in the initial system/control prefix | Early tokens change | Cache-breaking |
| Mistral Small 4 | Early MODEL_SETTINGS block | Early tokens change | Cache-breaking |
| Qwen3.8 | Graded effort becomes an initial system instruction | Early tokens change | Cache-breaking |
| gpt-oss | Opening Harmony system message | Early tokens change | Cache-breaking |
| MiniMax M3 | Thinking instruction in the initial system prompt | Early tokens change | Cache-breaking |
This snapshot covers the principal families investigated on 5 September 2026; it is not exhaustive. “Cache-breaking” describes the official serialization template, not an intrinsic limit of the weights.10131521223536373839
Qwen3 and Gemini show the boundary
The original Qwen3 follows the most recent /think or /no_think instruction, so a new user or system message can switch modes while prior tokens remain identical. A runtime such as vLLM can reuse matching prefix blocks. This is a real precursor, but a weaker contract: the control is binary, textual rather than privileged, vulnerable to imitation by user or retrieved content, and dependent on stable serialization, routing, eviction, and runtime configuration. There is no hosted guarantee equivalent to Astra's.932
Qwen3.8 illustrates the trade-off in the other direction. Its template turns graded effort into an instruction in the first system block. The control is finer, but changing it rewrites the start of the sequence and loses prefix composability.2235
Gemini separates a request's thinking level from an explicitly named cached-content object. An application can cache a large system prompt or document corpus, reference it repeatedly, and vary thinking per request. That preserves the economic value of stable RAG context, but Google does not promise that a changed thinking level preserves an arbitrary accumulated dialogue cache. Treat it as a near-equivalent and verify hits with usage.total_cached_tokens rather than inferring them from request structure.534
What changes for an agent harness
Cache-preserving transitions make adaptive effort practical at the step level. A harness can use low effort for extraction, routing, formatting, and routine tool calls; medium for normal execution; high for planning and ambiguous decisions; and the deepest level for recovery, consequential actions, or final review—without turning every policy change into a full-prefix miss.
The optimization unit is no longer just the model. It is the model plus prompt protocol, controller, cache implementation, and compaction strategy. Two applications using the same weights can have very different latency and cost because one maintains an immutable prefix while the other rewrites system instructions, tools, or effort on every turn.
Typed privileged controls are safer than textual switches such as /think: they are harder for users or retrieved documents to spoof, easier to validate and audit, and clearer training targets. This suggests an append-only inference control plane in which a harness can record events such as effort changes, tool grants, permission revocations, response verbosity, or deadlines without rewriting semantic history. That is an architectural inference from current APIs—not a provider-announced roadmap.
The unresolved problem is compaction
An append-only history eventually becomes too large. Compaction then rewrites the prefix and may erase the events that established current policy. A complete checkpoint has to preserve both summarized semantic history and effective runtime state: reasoning effort, active tools, permissions, output configuration, and outstanding agent state. Astra's current compaction restrictions expose the protocol gap directly. Future agent APIs will likely need explicit snapshots that combine compressed history with a typed configuration checkpoint.
Until then, a production harness should keep stable instructions and long-lived context first; append rather than rewrite; record cache reads, writes, and uncached tokens; track effective effort independently of response metadata; reapply state after compaction; and pin effort for a session when the model template encodes it at the beginning. Escalation is also forward-only: raising effort now cannot retroactively repair a bad decision already embedded in the history.
Choosing an effort level #
A reasonable starting rule is to use the lowest effort that reliably clears the quality bar. This is different from choosing the lowest effort that usually looks acceptable. The bar should be expressed as an observable outcome: a structured extraction validates, a bug diagnosis identifies the actual fault, a review finds known failure modes, or an agent completes a task without violating constraints.
Minimal or no reasoning is appropriate for latency-critical work that has a direct mapping from input to output. Classification, extraction, formatting, and simple routing often belong here. Low effort is useful when a task includes a small amount of judgment or a familiar tool call but still has one likely solution path.
Medium effort is a sensible experimental baseline for planning, debugging, and multi-step work. It gives the model room to organize a solution without immediately accepting the latency of the deepest settings. High effort becomes interesting when the task has several plausible approaches, interacting constraints, or errors that are expensive to miss. The highest settings should usually be reserved for difficult, quality-first work that can run asynchronously—and kept only when measured gains justify them.
A routing test
Before increasing effort, ask:
- Does the task require several dependent decisions?
- Are there multiple plausible approaches worth comparing?
- Would a subtle mistake be expensive or difficult to detect?
- Can the user or workflow tolerate additional latency?
- Do evaluations show that extra effort improves the outcome?
The first three questions estimate potential value. The fourth establishes the operational budget. The fifth decides the matter. When the evaluation says no, the higher setting is overhead, regardless of how sophisticated it sounds.
Task difficulty is not prompt length
Long inputs are not necessarily hard, and short inputs are not necessarily easy. A ten-page document may require a direct extraction. A two-sentence scheduling problem may contain a dense set of constraints. Useful routers look at the shape and stakes of the requested work rather than using token count as a difficulty score.
Evaluate a frontier, not a winner #
There is no globally best effort level. The goal is to find configurations on the quality–latency–cost frontier for a particular workload. One configuration dominates another when it is at least as good on every dimension and better on one. Everything else is a product decision about which trade-off matters.
What published work already shows
The most influential independent result is Snell and colleagues’ ICLR 2025 study of test-time compute. Using mathematical reasoning tasks, revision models, and process reward models, the authors found that the useful strategy depended on problem difficulty. Their adaptive, compute-optimal policy reached the performance of a best-of-N baseline with up to four times less inference compute. In a FLOPs-matched comparison, a smaller model with test-time compute could beat a model with fourteen times as many parameters on problems where the smaller model already had a non-trivial chance of success. On the hardest problems, however, additional inference produced little benefit; more capable weights remained the better investment.27
The peer-reviewed s1 project provides a concrete implementation case. The researchers fine-tuned Qwen2.5-32B-Instruct on 1,000 carefully selected reasoning examples, then controlled its inference budget by ending the reasoning span at a limit or appending “Wait” when the model tried to stop. On AIME 2024, extending the budget raised reported accuracy from 50% to 57%. The case is important because the data, model, and code are open, but its result should not be universalized: the model was specifically trained for the intervention and the headline evaluations were competition mathematics and science questions.28
A later large-scale study generated more than 30 billion tokens with eight open models across four reasoning datasets. It found no universally best test-time strategy. Some models favored short traces; others benefited from longer traces only on hard problems; expanding beam search often flattened or reduced accuracy even while consuming more tokens. The best policy depended on the model’s post-training, the problem, and the available compute budget.30
A published harness case: retry only the failures
Anthropic reported a directly operational experiment on an internal subset of SWE-bench Pro. Running Claude Opus 5 at low effort first, then rerunning only the test-detected failures at the default effort, achieved about a 93% pass rate for roughly $0.70 per task. Running every task once at the default achieved 91.7% for $1.39 per task. Starting at medium and rerunning failures produced about 94% for $0.95.29
The result is a useful case study of what an agent harness contributes: the model adapts within one call, but the harness observes an external test result and reallocates budget across calls. It also exposes the boundary conditions. The measurement was vendor-run on a private subset and is not comparable to the public SWE-bench leaderboard. It works because tests provide a reliable failure signal, and failed tasks pay for two attempts and therefore take longer.
What these results establish
- Additional inference compute can improve accuracy when the model has a viable solution path.
- The way compute is spent—longer trace, revisions, parallel samples, verifier, or retry—matters as much as the amount.
- The hardest tasks may need a stronger model rather than more effort on a weaker one.
- When outcomes are cheaply verifiable, low-first escalation can beat any fixed setting on cost per successful task.
These findings reject “always high” and “always low” as general policies. They give an engineer a defensible prior: allocate effort adaptively, use external verification when it is trustworthy, and measure success at the level of the complete task rather than one model call.
Routing effort in production #
Once evaluation shows that task families have different sweet spots, a single global default becomes wasteful. The production system can choose effort from observable request properties such as task type, number of constraints, tool plan, stakes, latency budget, or whether the work can run asynchronously.
function chooseEffort(task) {
if (task.kind === "extract" && task.schemaIsStrict) {
return "low";
}
if (task.isHighStakes || task.hasManyConstraints) {
return task.canRunAsync ? "high" : "medium";
}
return "medium";
}
Keep the router explainable and log its decision. Effort is otherwise an invisible variable during incident analysis: two apparently identical requests may have taken different paths for reasons nobody can reconstruct. A small set of named routing rules is often more useful than an opaque difficulty classifier, especially early in a product’s life.
The deepest settings belong naturally in asynchronous paths. Give them explicit token and wall-clock budgets, handle incomplete responses, and surface progress separately from the final answer. An interactive request should not silently turn into a multi-minute run because a router overestimated its difficulty.
Production guardrails
- Pin a model version while comparing configurations.
- Set generated-token and wall-clock limits.
- Handle incomplete responses before presenting output.
- Log the model, effort, usage, latency, router reason, and validator result.
- Re-run evaluations after model, prompt, tool, or schema changes.
Six ways teams waste money on reasoning #
1. Defaulting every request to high
This is operationally simple but usually inefficient. Routine work absorbs extra latency and cost, while difficult work still lacks a task-specific success criterion. A global high setting can conceal poor routing and weak evaluation rather than solve them.
2. Using answer length as a quality metric
Longer outputs often feel more considered, but visible detail and internal reasoning are different controls. Grade correctness, coverage, evidence, and downstream success. Otherwise an evaluation may reward eloquence rather than problem solving.
3. Changing the prompt and effort together
If a test rewrites the prompt, changes the tool set, and raises effort at the same time, it cannot show which change produced the result. Establish a baseline and sweep one variable at a time before testing interactions.
4. Ignoring incomplete responses
Check the response status. When reasoning consumes the output envelope, retry with a larger limit, lower effort, or a simpler task — don't present a truncated answer.
5. Treating labels as portable
”Medium” is not a standardized quantity. Different models can use different amounts of computation, adapt differently to the prompt, or support different levels entirely. Recalibrate whenever the model or provider changes.
6. Breaking prompt caches with the wrong kind of effort change
On models whose templates place effort near the beginning, changing it rewrites the cached prefix and forces the history to be processed again. Use a documented append-only transition where one exists; otherwise pin effort per session and measure total cached and uncached input, not only reasoning tokens per call.
The rules, for those who skim #
- Default to the lowest effort that clears the quality bar. Not the lowest that "usually looks fine" — the lowest that your evaluations say is reliable.
- Treat the setting as task-specific, not global. Classification and extraction at low. Multi-step reasoning at medium. Irreversible decisions at high. Never one level for everything.
- Start low, escalate on evidence. Run cheap, validate, retry at higher effort only on failure. This beats any fixed level on cost per successful task — when validation is reliable.
- Measure the frontier, not a winner. You're looking for configurations on the quality-latency-cost Pareto frontier for your workload. There is no globally best level.
- Log the native setting, not the gateway label. "Medium" is not portable across providers. Record model version, provider-native effort, token usage, latency, and validator result.
- Keep long-running histories append-only when the protocol allows it. Track cached and uncached input, and reapply the effective configuration after compaction.
- Higher effort can make things worse. The model may overcomplicate a direct question, hallucinate edge cases, or spend its budget on an unproductive line of reasoning. More is not always better.
- Re-evaluate after any change. New model version, new prompt, new tool set — recalibrate. Yesterday's sweet spot is today's guess.
Closing thoughts #
Buy reasoning where it changes the outcome. As effort becomes mutable runtime state, preserve cacheable history and treat the prompt protocol, controller, cache, and compaction strategy as one system.
The right default is not the setting that makes the model think the most. It is the lowest setting that makes the system work reliably.
Sources and further reading #
Provider controls
- OpenAI, “Reasoning models” — reasoning tokens, effort behavior,
configuration_update, cache preservation, compaction constraints, and usage semantics. - OpenAI, “Model guidance” — current GPT-5.6 effort levels and selection guidance.
- Anthropic, “Adaptive thinking” — per-request thinking decisions, effort steering, and interleaved thinking.
- Anthropic, “Effort” — response-wide and per-message effort, cache-preserving updates, supported levels, defaults, and model differences.
- Google, “Gemini thinking” — dynamic thinking, thinking levels, budgets, and model-specific support.
- OpenRouter, “Reasoning tokens” — normalized effort and token-budget controls plus provider mappings.
- Baseten, “DeepSeek V3.2” and Inference API overview — model-specific thinking through an OpenAI-compatible endpoint.
Open models and agent infrastructure
- DeepSeek, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning” — RLVR, R1-Zero, the multi-stage R1 pipeline, and distillation.
- Qwen, “Qwen3 inference with Transformers” — thinking and non-thinking modes, hard and soft switches, and trace parsing.
- OpenAI, “gpt-oss-120b” and “Harmony response format” — open weights, configurable effort, reasoning channels, and prompt protocol.
- OpenAI, “Agents SDK models” and “Usage” — agent-level reasoning settings and run-level token accounting.
- LangChain, “Agents” — middleware, dynamic model selection, state, retries, and tool orchestration.
- DeepSeek, “Thinking mode” and V4 encoding documentation — current effort levels, compatibility mappings, and prompt prefixes.
- NVIDIA, “Nemotron 3 Ultra” and its technical report — reasoning-budget control, post-training, and multi-teacher distillation.
- Moonshot AI, “Kimi K3” and its technical report — effort levels, preserved thinking, effort specialists, and on-policy distillation.
- Z.ai, “GLM-5.3” and “Thinking mode” — current effort API plus interleaved, preserved, and turn-level thinking.
- Thinking Machines Lab, “Inkling: Our Open-Weights Model” — continuous effort conditioning and per-token cost during large-scale RL.
Current release updates and synthesis
- Sebastian Raschka, “Controlling Reasoning Effort in LLMs” — a detailed coverage baseline for training recipes and open-weight implementations.
- Z.ai, “GLM-5.3-Flash: Frontier Intelligence, Flash Cost” — the latest Flash release and its measured behavior across effort levels.
- Moonshot AI, “Kimi K2.5: Visual Agentic Intelligence” — a multimodal agentic model report that includes the Token-Efficient RL (Toggle) method for controlling reasoning length through alternating budgeted and unconstrained training phases.
- xAI, “Reasoning” — Grok 4.6 effort levels, encrypted reasoning continuity, and the different semantics of effort on the multi-agent model.
- Qwen, “Qwen3.8” — the latest open-model series, explicit reasoning effort, and preserved thinking.
Historical foundations
- Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” — the 2022 result that established worked intermediate reasoning as an inference-time capability.
- Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models” — sampling diverse reasoning paths and selecting their most consistent answer.
- OpenAI, “Improving mathematical reasoning with process supervision” — rewarding intermediate steps and selecting among multiple solutions.
- OpenAI, “Learning to reason with LLMs” — the September 2024 o1 release, reinforcement learning, test-time compute, and hidden chain-of-thought.
Published evidence and case studies
- Snell et al., “Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters”, ICLR 2025 — adaptive allocation by problem difficulty, verifier search, revisions, and FLOPs-matched comparisons.
- Muennighoff et al., “s1: Simple Test-Time Scaling”, EMNLP 2025 — an open implementation of budget forcing trained from 1,000 curated reasoning examples.
- Anthropic, “Optimizing for cost and intelligence” — vendor-run effort sweeps, SWE-bench Pro escalation results, task budgets, and multi-model routing.
- Agarwal, Sengupta, and Chakraborty, “The Art of Scaling Test-Time Compute for Large Language Models” — a 30-billion-token comparison of test-time strategies across eight open models and four reasoning datasets.
Caching and mutable controls
- OpenAI, “GPT-6 Astra model guidance”, “Prompt caching”, and launch announcement — appended effort updates, reusable prefixes, cache semantics, and the 3 September 2026 release.
- vLLM, “Automatic Prefix Caching” — block-level reuse of matching token prefixes in open-model serving.
- Anthropic, “Release notes” — the 1 September 2026 per-message effort release and earlier mid-conversation system and tool changes.
- Google, “Context caching” — immutable explicit caches, implicit caching, cached-token accounting, and API boundaries.
- Qwen, Qwen3.8 chat template — graded effort translated into an instruction in the initial system block.
- Z.ai, GLM-5.3 chat template — reasoning effort serialized into the opening control and system prefix.
- Mistral AI, Mistral Small 4 chat template — reasoning effort in an early
MODEL_SETTINGSblock. - MiniMax, MiniMax M3 chat template — thinking-mode instructions in the initial system prompt.
- xAI, “Prompt caching for multi-turn conversations” — exact-prefix cache behavior without a documented guarantee across effort changes.