◈AI Codex
Foundation Models & LLMsHow It Works

Effort: how deeply should Claude think?

In brief

Extended thinking as a toggle is gone on current models — Claude Opus 5, Sonnet 5 and Fable 5 think by default and adaptively, and budget_tokens is not accepted. The control that replaced it is `effort`, which governs total token spend including tool calls. Here's when to raise it, when to lower it, and the two behaviours that will bite you.

5 min read·Extended Thinking

Contents

♡Sign in to save

Claude thinks before it answers. On Claude Opus 5, Claude Sonnet 5 and Claude Fable 5, that is not a mode you switch on — it is on by default, and the model decides for itself how much reasoning a given request deserves. This is adaptive thinking, and it replaced the older arrangement where you enabled "extended thinking" and handed the model a token budget.

If you learned this feature in its earlier form, the important correction is this: the manual configuration is gone on current models. thinking: {type: "enabled", budget_tokens: N} is not accepted on Claude Opus 5, Sonnet 5 or Fable 5. What you control now is effort.

The control that replaced the toggle

effort takes five values — low, medium, high, xhigh, max — and defaults to high. It does not turn thinking on or off. It tells Claude how much of everything to spend: reasoning, tool calls, and response tokens alike.

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=8192,
    messages=[{"role": "user", "content": "..."}],
    output_config={"effort": "medium"},
)

That last point is the one people miss. Effort is not a thinking dial, it is a total-spend dial. At low effort Claude makes fewer tool calls, combines operations, and skips the preamble. At high effort it explores more, calls more tools, and explains its plan. On agentic work this matters far more than the reasoning depth on any single turn.

Setting effort: "high" is identical to omitting the parameter.

When to raise effort above the default

Multi-step reasoning. Problems where you have to get step 3 right to get step 5 right — financial models, legal analysis, proofs, anything that chains dependencies.

Problems with many valid approaches. When you want Claude to weigh options against each other rather than commit to the first plausible one. Strategic decisions, architectural choices, competing interpretations.

Tasks where systematic coverage matters. Reviewing a contract for risk, auditing a plan for gaps, checking an argument for logical flaws — where missing something has a cost.

Long-horizon agentic and coding work. This is what xhigh exists for: tasks running over thirty minutes with token budgets in the millions, repeated tool calling, detailed search. On Opus 5 and Sonnet 5, xhigh is the recommended starting point for demanding coding work, not max.

max is narrower than it sounds. Reserve it for genuinely frontier problems. On most workloads it adds significant cost for small quality gains, and on structured-output tasks it can lead to overthinking.

When to lower it

High-volume or latency-sensitive work. Classification, extraction, routing, chat. low is the single most effective cost lever available on the current models, and on Fable 5 low effort still outperforms xhigh on earlier generations.

Subagents. A subagent doing one scoped mechanical job does not need the orchestrator's effort level. This is where teams leave the most money on the table.

Simple retrieval and reformatting. If the answer does not require reasoning, effort spent on reasoning is latency you are paying for and not using.

Creative and generative writing. More deliberation tends to make output more systematic and less fluent. For open-ended creative work, the reasoning process is as much a constraint as a capability.

The practical test

Run the same task at medium and at high against your actual inputs and compare. For work that genuinely benefits, the difference is obvious. For work that doesn't, you have just found a permanent cost reduction.

Do this per workload, not once. Effort levels do not transfer cleanly between models — if you carried settings over from an earlier model, run a fresh sweep rather than reusing them.

Two things that will bite you

Changing effort mid-conversation invalidates your prompt cache. Effort shapes the rendered prompt, so switching levels between requests means cached prefixes from earlier turns do not hit. Pick a level at the start of a session that relies on caching and hold it. Vary effort across workloads, not within one conversation.

Thinking cannot always be disabled. On Claude Opus 5, thinking: {type: "disabled"} works at high effort or below, but returns a 400 error at xhigh or max. Claude Fable 5 rejects it outright. Do not build a code path that assumes you can always turn thinking off.

Where the old configuration still applies

If you are calling an older model that supports only extended thinking, type: "enabled" with budget_tokens is still how you configure it — see the extended thinking documentation and the per-model configuration table. On Claude Opus 4.8, 4.7, 4.6 and Sonnet 4.6, thinking is adaptive but off until you set thinking: {type: "adaptive"}.

One more detail worth knowing: on the current models display defaults to "omitted", so thinking blocks come back empty unless you ask for them. Set thinking: {"type": "adaptive", "display": "summarized"} when you want to see the reasoning.

Further reading

  • Effort — the parameter, its levels, and per-model recommendations
  • Thinking — how thinking works and how it interacts with tools, caching, and streaming
  • Extended thinking — the older manual configuration, for models that still take it
  • Adaptive thinking — the mode current models run by default

Official training on this: Choosing the right effort level in Claude Code on Claude Academy, free.

Weekly brief

For people actually using Claude at work.

Each week: one thing Claude can do in your work that most people haven't figured out yet — plus the failure modes to avoid. No tutorials. No hype.

No spam. Unsubscribe anytime.

What to read next

All articles →