AIUnlimited
🌳

Fondations IA

đŸŒ±
AI Seeds

Partez de zéro

🌿
AI Sprouts

Construisez les fondations

🌳
AI Branches

Mettez en pratique

đŸ•ïž
AI Canopy

Approfondissez

đŸŒČ
AI Forest

MaĂźtrisez l'IA

🔹

MaĂźtrise IA

✏
AI Sketch

Partez de zéro

đŸȘš
AI Chisel

Construisez les fondations

⚒
AI Craft

Mettez en pratique

💎
AI Polish

Approfondissez

🏆
AI Masterpiece

MaĂźtrisez l'IA

📘

Pratique IA

📖
Comprendre les modĂšles open-source

Fondamentaux et ressources pour les modĂšles open-source

🎯
Du problĂšme Ă  la tĂąche modĂšles

Transformer les problÚmes métier en tùches modÚles

⚡
Exécuter votre premier modÚle

Voyez vos premiers résultats en 30 minutes

🔧
Fine-tuning et évaluation

Affinez les modÚles et évaluez les performances

🚀
SystĂšmes d'application

Construisez des applications IA réelles

🎹
IA générative

Explorez les modĂšles AIGC open-source

đŸ€–
Agents

Apprenez les frameworks Agent et outils MCP

📐
Fondamentaux supplémentaires

Bases LLM et évaluation

🎓

Claude Académie

đŸ€–
Claude 101

Learn AI basics with Claude

đŸ’»
Claude Code 101

Code with Claude as your pair programmer

đŸ€
Introduction to Claude Cowork

Collaborate with Claude on complex projects

⚙
Claude Platform 101

Build apps with the Claude API

Labo

7 expériences chargées
🧬Bac Ă  sable neuronalđŸ€–IA ou Humain ?đŸ„‹Dojo d'ingĂ©nierie de prompts🏁Course d'algorithmes🧠Quiz IAđŸ—ïžCanevas de conception systĂšme
🎯Entretien simulĂ©Entrer dans le labo→
🚀

Développement de carriÚre

🚀
Rampe de lancement entretien

Commencez votre parcours

🌟
MaĂźtrise comportementale

Maßtrisez les compétences relationnelles

đŸ’»
Entretiens techniques

Réussissez l'épreuve de code

đŸ€–
Entretiens IA et ML

MaĂźtrisez l'entretien ML

🏆
Offre et au-delĂ 

Décrochez la meilleure offre

Commencer
AIUnlimited

Licence MIT

æČȘICP怇18025655ć·-11

Apprendre

  • Bases de l'IA
  • Pratique IA
  • Claude AcadĂ©mie
  • Labo
  • DĂ©veloppement de carriĂšre

Communauté

  • À propos
  • FAQ

Soutien

  • Conditions d'utilisation
  • Politique de confidentialitĂ©
  • Contact
Programmes d'IA et d'ingĂ©nierieâ€ș⚙ Claude Platform 101â€șLeçonsâ€șChoosing the Right Model
🎯
Claude Platform 101 ‱ IntermĂ©diaire⏱ 5 min de lecture

Choosing the Right Model

Choosing the right model

You're shipping an app with Claude. Which model do you pick? If you default to the smartest one, your API bill will surprise you. Pick the cheapest one, and the output might not hold up. Each model has different trade-offs, and picking the right one affects both quality and cost.

The model tiers

Anthropic currently offers four model tiers, and you choose between them with the model parameter in your API call.

Note that Claude Fable 5.1 has been generally available since September 1, 2026, but is not reflected in the video above. Learn more about Claude Fable 5.1 and Claude Mythos 5.1 here. The video and terminal screenshot in this lesson were recorded with earlier models (Claude Opus 4.7 and Claude Sonnet 4.6). The code below uses the current model IDs; your latency and token numbers will differ.

  • Claude Fable is our most capable model yet — a new tier that sits above Opus, built for your toughest challenges. It comes at a higher cost than Opus, so reserve it for work where that extra capability is worth paying for. The current Fable model is Claude Fable 5.1 (claude-fable-5-1).
  • Claude Opus is the most capable of the three core model families, but also the slowest and highest cost of the three. Use it for deep reasoning, complex analysis, multi-step coding, and nuanced writing. The current Opus model is Claude Opus 5 (claude-opus-5).
  • Claude Haiku is the fastest and lowest cost, optimized for speed and cost efficiency rather than maximum intelligence. Use it for high-volume, low-complexity work like classification, extraction, and routing. The current Haiku model is Claude Haiku 4.5 (claude-haiku-4-5).
  • Claude Sonnet sits in the sweet spot: a balanced combination of intelligence, speed, and cost that works well for most production work. The current Sonnet model is Claude Sonnet 5 (claude-sonnet-5).
Leçon 3 sur 130% terminé
←Your First API Call

Discussion

Se connecter pour rejoindre la discussion

Start with a simple evaluation

Before you write production code, set up a simple evaluation: a set of example inputs that you run through each model and score against what good output means for your use case. You don't need anything fancy — 20 or 30 representative examples from your actual workload is enough to start.

Then work your way up the tiers:

  1. Run your examples through Haiku first. If the quality holds, you're done — and you just saved a lot of money.
  2. If it doesn't, step up to Sonnet.
  3. Only reach for Opus when the task needs it.

Comparing the tiers side by side

Let's see the difference between the tiers, not just talk about it. We'll send the same prompt through all three models and watch the latency and token counts:

models = ["claude-haiku-4-5", "claude-sonnet-5", "claude-opus-5"]

for model in models:
    response = client.messages.create(
        model=model,
        max_tokens=300,
        messages=[{"role": "user", "content": prompt}],
    )
    print(model, response.usage)

Two things are going on here:

  • The loop swaps the model field on each request. Same prompt, same max tokens — only the model changes.
  • response.usage gives you the input and output tokens straight back from the API, which is what your bill is calculated on.

Run it and you'll see three models and three sets of numbers. Opus takes the longest and reads the most polished — but for a two-sentence definition, that polish is wasted. Sonnet tightens the writing up a little. And Haiku comes back, often in under a second, with a very competent two-sentence answer. It's honestly perfect for this kind of scenario.

And that's the whole point: the right model is the cheapest one whose output you'd actually ship. For a definition, Haiku is plenty. For drafting a regulatory response, you'd run the same comparison and probably end up on Opus. The eval is the same shape every single time.

Routing different work to different models

In a real app, you'd route different kinds of work to different models inside the same endpoint. Take an operations dashboard with a document processing route:

  • Every incoming file gets classified with Haiku.
  • Client updates get drafted with Sonnet.
  • Only RFP responses reach for Opus.

One queue, three models, picked per task.

Recap

  • Anthropic offers four model tiers: Fable for the highest available capability, Opus for hard problems, Sonnet for daily work, and Haiku for volume.
  • Set up a simple evaluation — 20 or 30 representative examples from your real workload — before writing production code.
  • Run the eval from Haiku upward and stop at the cheapest model whose output you'd actually ship.
  • response.usage reports input and output tokens, which is what your bill is based on.
  • In production, route different tasks to different models inside the same endpoint instead of picking one model for everything.