Some tasks need more than a quick answer. Claude can work through a problem before responding â a feature called extended thinking. In this lesson, we'll look at what thinking is, how it works, and when it actually helps.
Here's the failure mode we're trying to avoid. Ask a model a multi-step question and have it answer immediately, and it can confidently get it wrong.
Extended thinking lets Claude reason step by step before producing a final response. When it's enabled, Claude generates internal reasoning tokens â often called a chain of thought â and then delivers the answer. The reasoning isn't hidden: you can see it in the response alongside the final text.
On Opus 5, thinking is adaptive and on by default. There's no token budget to pick: Claude decides dynamically when to think and how much.
To control how much Claude thinks, use the effort parameter. One gotcha: it goes inside output_config, not next to the thinking block. The levels are:
lowmediumhigh (the default)xhigh (extra high)maxExtended thinking helps with:
Skip it for simple classification, extraction, or boilerplate. For those tasks it just adds latency and cost without actually improving the results.
Let's see it work. Here's an agent loop with one weather tool, and we'll ask Claude to plan a road trip out of San Francisco â two stops, weighing weather and drive time. That's a real trade-off, the kind of question where thinking earns its keep.
Se connecter pour rejoindre la discussion
import anthropic
client = anthropic.Anthropic()
weather_tool = {
"name": "get_weather",
"description": "Get the current weather for a city.",
"input_schema": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"}
},
"required": ["city"],
},
}
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
thinking={"type": "adaptive", "display": "summarized"}, # summarized = return the reasoning text
output_config={"effort": "high"}, # low | medium | high | xhigh | max
tools=[weather_tool],
messages=[
{
"role": "user",
"content": "Plan a road trip out of San Francisco with two stops, "
"weighing weather and drive time.",
}
],
)
When you run this, the output is more interesting than usual. You'll see thinking blocks where Claude works through the trade-offs, followed by tool calls to check each city, and finally a text block with the actual recommendation.
The reasoning is visible â that's the whole point.
In a production app, this is the difference between an agent that finds problems one at a time and an agent that connects them. Take a compliance review app: toggling adaptive thinking on the auto-review call lets the agent reason across report sections â catching things like a wind load spec in section three that conflicts with the material spec elsewhere in the document.
"display": "summarized" to see the reasoning in the response.output_config: low, medium, high (default), xhigh, or max.