You're shipping an app with Claude. Which model do you pick? If you default to the smartest one, your API bill will surprise you. Pick the cheapest one, and the output might not hold up. Each model has different trade-offs, and picking the right one affects both quality and cost.
Anthropic currently offers four model tiers, and you choose between them with the model parameter in your API call.
Note that Claude Fable 5.1 has been generally available since September 1, 2026, but is not reflected in the video above. Learn more about Claude Fable 5.1 and Claude Mythos 5.1 here. The video and terminal screenshot in this lesson were recorded with earlier models (Claude Opus 4.7 and Claude Sonnet 4.6). The code below uses the current model IDs; your latency and token numbers will differ.
claude-fable-5-1).claude-opus-5).claude-haiku-4-5).claude-sonnet-5).साइन इन करें चर्चा में शामिल हों
Before you write production code, set up a simple evaluation: a set of example inputs that you run through each model and score against what good output means for your use case. You don't need anything fancy — 20 or 30 representative examples from your actual workload is enough to start.
Then work your way up the tiers:
Let's see the difference between the tiers, not just talk about it. We'll send the same prompt through all three models and watch the latency and token counts:
models = ["claude-haiku-4-5", "claude-sonnet-5", "claude-opus-5"]
for model in models:
response = client.messages.create(
model=model,
max_tokens=300,
messages=[{"role": "user", "content": prompt}],
)
print(model, response.usage)
Two things are going on here:
model field on each request. Same prompt, same max tokens — only the model changes.response.usage gives you the input and output tokens straight back from the API, which is what your bill is calculated on.Run it and you'll see three models and three sets of numbers. Opus takes the longest and reads the most polished — but for a two-sentence definition, that polish is wasted. Sonnet tightens the writing up a little. And Haiku comes back, often in under a second, with a very competent two-sentence answer. It's honestly perfect for this kind of scenario.
And that's the whole point: the right model is the cheapest one whose output you'd actually ship. For a definition, Haiku is plenty. For drafting a regulatory response, you'd run the same comparison and probably end up on Opus. The eval is the same shape every single time.
In a real app, you'd route different kinds of work to different models inside the same endpoint. Take an operations dashboard with a document processing route:
One queue, three models, picked per task.
response.usage reports input and output tokens, which is what your bill is based on.