AIUnlimited
🌳

Fondations IA

đŸŒ±
AI Seeds

Partez de zéro

🌿
AI Sprouts

Construisez les fondations

🌳
AI Branches

Mettez en pratique

đŸ•ïž
AI Canopy

Approfondissez

đŸŒČ
AI Forest

MaĂźtrisez l'IA

🔹

MaĂźtrise IA

✏
AI Sketch

Partez de zéro

đŸȘš
AI Chisel

Construisez les fondations

⚒
AI Craft

Mettez en pratique

💎
AI Polish

Approfondissez

🏆
AI Masterpiece

MaĂźtrisez l'IA

📘

Pratique IA

📖
Comprendre les modĂšles open-source

Fondamentaux et ressources pour les modĂšles open-source

🎯
Du problĂšme Ă  la tĂąche modĂšles

Transformer les problÚmes métier en tùches modÚles

⚡
Exécuter votre premier modÚle

Voyez vos premiers résultats en 30 minutes

🔧
Fine-tuning et évaluation

Affinez les modÚles et évaluez les performances

🚀
SystĂšmes d'application

Construisez des applications IA réelles

🎹
IA générative

Explorez les modĂšles AIGC open-source

đŸ€–
Agents

Apprenez les frameworks Agent et outils MCP

📐
Fondamentaux supplémentaires

Bases LLM et évaluation

🎓

Claude Académie

đŸ€–
Claude 101

Learn AI basics with Claude

đŸ’»
Claude Code 101

Code with Claude as your pair programmer

đŸ€
Introduction to Claude Cowork

Collaborate with Claude on complex projects

⚙
Claude Platform 101

Build apps with the Claude API

Labo

7 expériences chargées
🧬Bac Ă  sable neuronalđŸ€–IA ou Humain ?đŸ„‹Dojo d'ingĂ©nierie de prompts🏁Course d'algorithmes🧠Quiz IAđŸ—ïžCanevas de conception systĂšme
🎯Entretien simulĂ©Entrer dans le labo→
🚀

Développement de carriÚre

🚀
Rampe de lancement entretien

Commencez votre parcours

🌟
MaĂźtrise comportementale

Maßtrisez les compétences relationnelles

đŸ’»
Entretiens techniques

Réussissez l'épreuve de code

đŸ€–
Entretiens IA et ML

MaĂźtrisez l'entretien ML

🏆
Offre et au-delĂ 

Décrochez la meilleure offre

Commencer
AIUnlimited

Licence MIT

æČȘICP怇18025655ć·-11

Apprendre

  • Bases de l'IA
  • Pratique IA
  • Claude AcadĂ©mie
  • Labo
  • DĂ©veloppement de carriĂšre

Communauté

  • À propos
  • FAQ

Soutien

  • Conditions d'utilisation
  • Politique de confidentialitĂ©
  • Contact
Programmes d'IA et d'ingĂ©nierieâ€ș⚙ Claude Platform 101â€șLeçonsâ€șContext Management
📩
Claude Platform 101 ‱ AvancĂ©â±ïž 6 min de lecture

Context Management

Context management

Every request you send Claude has a context window. A million tokens sounds like a lot, but it runs out faster than you think once you're shipping a real agent. That's where context management comes in: it's how you stay inside the window without losing what matters.

What counts as context

Context is everything Claude sees on a given turn:

  • The system prompt
  • The message history
  • Tool definitions and tool results
  • Attached files and skills
  • Thinking blocks

It's the input to every single API call. You pay for it on the way in, and you pay for it on the way out. And once the window is full, the request fails.

So the goal isn't to fit everything in. The goal is to fit the right things in.

Anthropic publishes four patterns for managing context in long-running agents. Three are first-class API features, and one is a design pattern.

Pattern 1: Just-in-time context

Don't load everything upfront. Load what the agent needs now, and let it pull more in via tools when it asks.

Think of a compliance review agent. It doesn't get the entire building code book stuffed into its system prompt — it calls a lookup_building_code tool when it needs a specific section. This is the design pattern of the four: nothing special in the API, just a deliberate choice about what you load and when.

Pattern 2: Server-side compaction

When a conversation runs long, Anthropic's server-side compaction summarizes old turns into a single block. You opt in by adding a context_management key to your request, holding an edit with a type:

response = client.beta.messages.create(
    betas=["compact-2026-01-12"],
    model="claude-opus-5",
    max_tokens=1024,
    context_management={
        "edits": [
            {"type": "compact_20260112"}
        ]
    },
    messages=messages,
)

The API auto-summarizes when the input crosses the trigger threshold. You don't have to track conversation length yourself.

Pattern 3: Prompt caching

Prompt caching lets you mark the stable parts of a request — the system prompt, the tool definitions, a long document — and reuse them across calls at a fraction of the cost.

Leçon 10 sur 130% terminé
←MCP

Discussion

Se connecter pour rejoindre la discussion

The math matters more than it looks. If your system prompt is 4,000 tokens and you call it 100 times an hour, caching is the difference between a usable bill and a phone call from finance.

Pattern 4: The memory tool

Some context needs to survive across sessions: user preferences, the agent's running notes, what was decided last week. The recommended primitive for this is the memory tool.

Here's how it works:

  • Claude reads and writes to a memory directory via tool calls.
  • You implement the storage backend client-side — a file system, a database, an encrypted store, whatever you want.
  • Anthropic auto-injects a system instruction telling Claude to check the memory directory before starting work.

Layering the patterns

In a production app, you'll usually layer all four at once. The compliance review agent caches its system prompt and tool definitions, and pulls building code sections in just in time via lookup_building_code.

Each pattern handles a different failure mode: cost, window size, statelessness. Pick the ones that match what's breaking for you.

Recap

  • Context is everything Claude sees on a turn — and it isn't free or infinite. Once the window fills, the request fails.
  • Just-in-time context: load what's needed now, let tools pull in the rest. This is the design pattern of the four.
  • Server-side compaction: add a context_management key, and the API summarizes old turns automatically when input crosses the trigger threshold.
  • Prompt caching: mark stable parts of the request and reuse them across calls at a fraction of the cost.
  • The memory tool: Claude reads and writes a memory directory via tool calls; you own the storage backend, so context survives across sessions.
  • Four patterns, one goal. Wire them up by hand, or use Claude managed agents, which ship with caching and compaction on by default.