◈AI Codex
Foundation Models & LLMsHow It Works

The context window in practice: what it means for how you work

In brief

The context window shapes what Claude can and can't do in any given conversation. Here is how to work with it.

5 min read·Context Window

Contents

♡Sign in to save

The context window is the amount of text Claude can hold in its working memory at once — everything from the current conversation, uploaded documents, your Project instructions, and its own responses. Claude Fable 5, Claude Opus 5 and Claude Sonnet 5 each have a 1,000,000 token context window — roughly 555,000 words, or several thousand pages. Claude Haiku 4.5 has 200,000. Maximum output is 128,000 tokens on the current large models, 64,000 on Haiku.

That sounds like more room than anyone could use, and for most work it is. But a big window changed which problem you have rather than removing it: you will rarely run out of space now, and you can still degrade your own output by filling that space with material Claude did not need.

What goes into the context window

Every token in the context window costs something, and affects how Claude attends to different parts of the conversation. The context window contains, in order:

  1. Your system prompt or Project instructions
  2. Any documents you have uploaded to the Project
  3. The full conversation history — every message from you and every response from Claude
  4. Your current message

The more that is in the context window, the more Claude has to process — and the more it may weight earlier instructions less heavily as the conversation gets longer.

The practical implications

Long conversations drift. In a very long conversation (40+ exchanges), Claude may start to lose track of instructions given early in the conversation, or produce outputs that are less consistent with the original setup. This is a property of attention — the model processes the full context, but recent content gets more weight. For long, complex tasks, it is often better to start a new conversation with a fresh context than to continue an old one indefinitely.

Big documents no longer threaten the limit, but they still cost you. A 50-page PDF is 30,000-plus tokens. Against a million-token window, thirty of those still fit. What they cost is money and latency on every turn, and attention: the more competing material in the window, the more the thing you actually care about has to compete with. Be selective because it produces better answers, not because you are about to run out.

Project instructions are always present. Your Project system prompt is sent with every message. If it is 3,000 tokens, that is 3,000 tokens consumed on every turn. This is why keeping Project instructions tight matters — see the guidance on writing system prompts for how to do this well.

Prompt caching helps at scale. For teams using Claude via the API, prompt caching lets you cache repeated context (like a large document that is referenced in every call) so it does not need to be re-processed each time. This is a significant cost and latency saving for high-volume use cases.

When context length actually matters

For most everyday use — drafting emails, answering questions, producing reports — you will never come close to the limit. A million tokens is more than almost any single task needs.

Context management matters when you are:

  • Working with document sets in the hundreds of pages
  • Running long analytical conversations where early instructions need to keep holding
  • Building workflows that involve many back-and-forth exchanges
  • Using the API for high-volume automated tasks, where every token in the window is billed on every turn

Note that very large contexts can carry their own pricing tier — check the pricing page before you design a workflow that routinely runs near the top of the window.

For Claude.ai users: if a conversation is getting long and Claude seems to be losing track of earlier instructions, start a fresh conversation. Your Project instructions reload from scratch, and you get clean attention.

The quality principle

A smaller context with the right information in it produces better outputs than a large context full of noise. Claude attends better to focused, relevant input. Before uploading a document, ask: does Claude actually need all of this? Often the answer is no — a 5-page summary of a 50-page document does better than the full document for most tasks.

This is the same principle as a good briefing: give someone exactly what they need to do the work, not everything you have.

Further reading

Official training on this: AI Capabilities and Limitations (13 lessons · 3.5 hr) · Parametric memory and context on Claude Academy, free.

Weekly brief

For people actually using Claude at work.

Each week: one thing Claude can do in your work that most people haven't figured out yet — plus the failure modes to avoid. No tutorials. No hype.

No spam. Unsubscribe anytime.

What to read next

All articles →