Claude Code context window: sizes, /context and compaction
Updated 8 min read
Claude Code’s context window is 1 million tokens on current models (Fable 5.1 and 5, Opus 4.7 and later, Sonnet 5 and later) on every paid plan, including Pro, with nothing to switch on. Opus 4.6 and Sonnet 4.6 run at 200K unless you select their [1m] variant, which depends on your plan. Everything in a session shares that space: the system prompt, CLAUDE.md, memory, MCP tool names, every file Claude reads and the conversation itself. Run /context to see what is using it; Claude Code compacts the conversation automatically as it fills.
Sizes below come from Anthropic’s model configuration docs and help center, read on 1 October 2026. Flags come from claude --help on Claude Code 2.1.274, the version on the machine we wrote this on.
How big is Claude Code’s context window?
| Model | Window in Claude Code | What you need |
|---|---|---|
| Fable 5.1, Fable 5 | 1M | Any paid plan, nothing to select |
| Opus 5.5, Opus 5, Opus 4.8, Opus 4.7 | 1M | Any paid plan, nothing to select |
| Sonnet 5.5, Sonnet 5 | 1M (there is no 200K version) | Any plan, no usage credits |
| Opus 4.6 | 200K, or 1M as claude-opus-4-6[1m] |
1M is included on Max, Team and Enterprise; Pro needs usage credits |
| Sonnet 4.6 | 200K, or 1M as claude-sonnet-4-6[1m] |
1M needs usage credits on every subscription, including Max (usage-based Enterprise excepted) |
On an API key or pay-as-you-go, both [1m] variants are fully available. Select one with /model:
/model opus[1m]
/model claude-sonnet-4-6[1m]
Details other pages leave out:
- Claude Code gets more than chat. The help center lists Opus 4.8, 4.7 and 4.6, Sonnet 4.6 and Fable 5 at 500K tokens in chat. The same models get 1M in Claude Code.
- No long-context premium. The 1M window “uses standard model pricing with no premium for tokens beyond 200K.”
- Cloud providers and gateways can be smaller. On Amazon Bedrock, Google Cloud’s Agent Platform and Microsoft Foundry, Opus 4.8 and later can run at 200K. Behind a gateway that stops at 200K, run
/autocompact 200k, since Claude Code cannot detect the lower limit. - Capping it.
CLAUDE_CODE_DISABLE_1M_CONTEXT=1holds every model to 200K. - Version matters. Opus 5.5 needs Claude Code 2.1.280 or later and Sonnet 5.5 needs 2.1.284. Our machine was on 2.1.274, so check
claude --versionfirst.
What fills the context window?
Anthropic’s context window page plays a sample session with representative token counts. Before you type anything, this loads:
| What loads at startup | Tokens in Anthropic’s example | Shown in your terminal? |
|---|---|---|
| System prompt | 4,200 | No |
Auto memory (MEMORY.md, first 200 lines or 25KB) |
680 | No |
| Environment info (directory, platform, git) | 280 | No |
| MCP tool names (full schemas deferred) | 120 | No |
| Skill descriptions, one line each | 450 | No |
~/.claude/CLAUDE.md |
320 | No |
Project CLAUDE.md |
1,800 | No |
That is 7,850 tokens before your first prompt: about 4% of a 200K window, under 1% of 1M. An AGENTS.md or output style can add more.
Then the work starts, and most of it is also invisible:
- File reads. You see “Read auth.ts”; Claude gets all 2,400 tokens. The docs say file reads “dominate context usage.”
- Path-scoped rules in
.claude/rules/load when Claude reads a matching file. - Command output. You see the test pass count; Claude gets 1,200 tokens of output.
- Hooks. A PostToolUse hook reaches Claude only through
additionalContext; plain stdout on exit 0 goes to the debug log.
How do I check my context usage in Claude Code?
/context
/context all
/context draws your usage as a colored grid by category, lists which CLAUDE.md and auto memory files loaded, and suggests fixes for context-heavy tools and memory bloat. Over the limit, it says by how much and which command frees space. In fullscreen mode, /context all expands the per-item list.
/usage shows session cost and plan limits, and on Pro, Max, Team and Enterprise flags anything behind 10% or more of recent usage, such as long context. /memory opens your CLAUDE.md and memory files, and /status shows the current model.
One way to measure your own overhead, which we have not tried: claude --help describes --safe-mode as starting with CLAUDE.md, skills, plugins, hooks and MCP servers disabled. Compare /context there with a normal session.
What happens when the context window fills up?
Claude Code compacts automatically: it clears older tool output first, then summarizes the conversation. The docs say your requests and key code snippets are kept, while “detailed instructions from early in the conversation may be lost.”
When it runs, by default:
- Models with a native 1M window: at about 967K tokens.
- Opus 4.6 and Sonnet 4.6 without
[1m]: at 200K, as for Opus 4.8 and later where they run at 200K. - Anything else: at the model’s limit. Cloud sessions compact as they approach it.
You can make it run earlier. /autocompact saves the value to your user settings, the --autocompact flag applies to one launch, and the environment variable overrides both:
# inside a session
/autocompact 500k
/autocompact auto
# one launch
claude --autocompact 300k
# scripts and CI: plain token count only
CLAUDE_CODE_AUTO_COMPACT_WINDOW=400000 claude
Accepted sizes run from 100K to 1M tokens, capped at the model’s window. If one huge output refills the window after each summary, Claude Code stops after a few attempts with a thrashing error.
What survives /compact?
| Content | After compaction |
|---|---|
| System prompt and output style | Still apply |
Project-root CLAUDE.md, rules without paths:, auto memory |
Re-read from disk |
Rules with paths:, nested CLAUDE.md files |
Summarized away; reload when Claude reads a matching file |
| Files Claude read or edited | Up to five re-read, most recently modified first; a file over 5,000 tokens comes back as a path only |
| Skills you invoked | Re-injected, up to 5,000 tokens each and 25,000 total, oldest dropped first |
| The skill list | Not re-injected |
| Context hooks added | Summarized with the rest |
So a rule that must survive belongs in the root CLAUDE.md without paths:, and a skill’s key instructions go at the top of SKILL.md, since truncation keeps the start.
Steer the summary with /compact focus on the auth bug fix, a “Compact instructions” section in CLAUDE.md (example in our context engineering guide), or /rewind and “Summarize from here”. /compact reads the whole conversation, so on a large context it is a large request; /clear costs nothing.
How do I keep a Claude Code session lean?
A 1M window is no reason to fill it. Claude Code sends the full conversation with every request, so a one-line question in a session open all day draws usage for all of it, at the cached rate. And Anthropic’s best practices say “performance degrades as it fills.”
- Clear between tasks.
/renamethe session,/clear, and/resumeit later./clearkeeps project memory. - Name the file. “Fix the bug in auth.ts” reads fewer files than “fix the login bug.”
- Send research to a subagent. In Anthropic’s example a subagent read 6,100 tokens of files and returned 420.
- Keep CLAUDE.md under 200 lines. Move procedures to skills and area rules to
.claude/rules/withpaths:.@imports do not save context. Block-level HTML comments are stripped before loading, so notes for humans cost nothing. More in what to put in CLAUDE.md. - Leave MCP tool search on. Only tool names load until a tool is used;
ENABLE_TOOL_SEARCH=falseloads every schema. Turn off unused servers in/mcp(see Claude Code MCP). - Hide skills with side effects.
disable-model-invocation: truekeeps a skill out of context until you type/name. - Filter noisy output. A PreToolUse hook can rewrite test commands to print only failures; the costs page has the script.
- Ask side questions with
/btw. They do not join the conversation history.
Where should long-term knowledge go?
Not in the window. A compaction summary lasts one session, and /clear drops it. Claude Code’s own answer is two files that reload from disk: CLAUDE.md, which you write, and auto memory in ~/.claude/projects/<project>/memory/, which Claude writes. Only the first 200 lines or 25KB of MEMORY.md load; topic files beside it are read when needed. Our Claude Code memory guide covers both.
The pattern is a short index that always loads and detail fetched on demand. It works until a second agent needs the same knowledge: auto memory is per repository, lives in one home directory, and other agents such as Codex or Cursor do not load it.
Share it with your other agents
Dexio is a hosted wiki that AI agents read and write through MCP. Claude Code, Codex, Cursor, Hermes Agent, OpenClaw, Claude, ChatGPT and other MCP clients connect to one wiki of linked markdown pages, and you see what your agents know as a page graph.
For the context window, that means:
- Little at startup. With tool search on, only Dexio’s tool names and server instructions load.
- Pages on demand. Claude finds a page with
search_pagesand loads it withread_page, which can read one section instead of the whole page. - Knowledge outlives the session. What Claude writes survives
/clearand compaction, and your other agents read the same page.
Point CLAUDE.md at the wiki in two lines instead of copying knowledge into it:
## Shared wiki
- Before working on a topic, search the Dexio wiki and read only the page or section you need.
- Record decisions and findings on the wiki page that covers them, not in this file.
Setup is in the Claude Code memory guide. Dexio is open source (AGPL-3.0) and free for one person; Team is $10 and Business $20 a member a month. To see a folder of notes as a graph first, try the wiki graph viewer, and for the wider picture see agent memory.
FAQ
Is Claude Code’s context window 200K or 1M? It depends on the model. Current Fable, Opus (4.7 and later) and Sonnet (5 and later) models get 1M on every paid plan. Opus 4.6 and Sonnet 4.6 get 200K unless you select [1m].
Does /clear delete CLAUDE.md or auto memory? No. It starts a new conversation with empty context and keeps project memory.
Why did Claude forget an instruction after compacting? Instructions given early in chat can be lost in the summary, and path-scoped rules are summarized away until a matching file is read again. Put rules that must last in the root CLAUDE.md.
Sources
- https://code.claude.com/docs/en/context-window
- https://code.claude.com/docs/en/model-config
- https://code.claude.com/docs/en/costs
- https://code.claude.com/docs/en/memory
- https://code.claude.com/docs/en/commands
- https://code.claude.com/docs/en/how-claude-code-works
- https://code.claude.com/docs/en/best-practices
- https://support.claude.com/en/articles/8606394-how-large-is-the-context-window-on-paid-claude-plans
- https://dexio.wiki/agents.md
- https://dexio.wiki/docs/
- https://dexio.wiki/pricing/