When to raise Claude's effort — and when it's a waste
In brief
You no longer turn thinking on — Claude Opus 5, Sonnet 5 and Fable 5 think by default. The question now is whether you're paying for more deliberation than the task returns. How to set `effort` per workload, and why leaving everything at the default is the most common avoidable cost.
Contents
Thinking used to be something you turned on. On the current models it isn't: Claude Opus 5, Claude Sonnet 5 and Claude Fable 5 think by default, and decide for themselves how deeply based on what you asked. The old budget_tokens configuration is not accepted on any of them.
What you steer instead is effort — five levels, low through max, defaulting to high. And the question worth asking is no longer "should I turn thinking on for this," because it is already on. It is "am I paying for more deliberation than this task returns."
For most operators, the answer is yes, and nobody checks.
Why the default is not always the right level
high is a sensible default for hard work. It is a poor default for volume. Effort governs everything Claude spends — reasoning, tool calls, response length — so leaving every workload at high means routine classification and extraction jobs are running at the setting designed for difficult coding problems.
The teams who notice this find their largest cost reduction here, and it costs them nothing in quality, because the work never needed the depth.
Raise effort when the problem has hidden complexity
Multi-step reasoning. Problems where you have to get step 3 right to get step 5 right. Financial models, legal analysis, anything that chains dependencies.
Ambiguous or contradictory inputs. When what you have given Claude is incomplete or internally inconsistent, more deliberation makes it surface that explicitly rather than gloss over it.
High-stakes decisions. When the cost of a wrong answer justifies the extra spend. Strategic analysis, risk assessments, anything you will act on directly.
Inconsistent results. If Claude keeps giving different answers to the same question, higher effort usually produces more stable, considered output.
Long agentic runs. xhigh exists for work measured in tens of minutes and millions of tokens — repeated tool calling, detailed search, coding across many files.
Lower it for routine work
First drafts, reformatting, summarising a document, answering a factual question, classification at volume, and any subagent doing one scoped job. low and medium are not degraded modes — they are the correct setting for work that does not need deliberation, and they are meaningfully faster.
Think of it as the difference between asking a colleague a quick question and booking an hour to work through a problem together. Both are right in different situations. Using the meeting format for every question is just expensive.
A practical rule
If you would be comfortable reading only the final answer without checking how Claude got there, you are probably paying for effort you do not need. If you would want to check the reasoning — because the problem is hard enough that the path matters — the default or higher is right.
Set it per workload and measure. The goal is appropriate depth, not maximum depth.
Further reading
- Effort — levels, defaults, and per-model guidance
- Thinking — what thinking does and when it runs
- Claude cost optimization — where effort fits among the other levers