Context engineering: a practical guide for AI agents
Updated 8 min read
Context engineering is deciding what goes into a model’s context window at each step: the instructions, tools, retrieved knowledge, memory and history an agent sees when it acts. Anthropic’s rule is to find “the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.” In practice that means short instruction files, lean tools, knowledge loaded when needed, and history that gets compacted or cleared. When you run more than one agent, it also means one place they all read and write, kept current.
What is context engineering?
Context is every token the model sees when it generates a response. Anthropic’s Applied AI team, in Effective context engineering for AI agents (September 2025), defines context engineering as “the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference, including all the other information that may land there outside of the prompts.”
Andrej Karpathy’s shorter version, quoted in LangChain’s post, calls it the “delicate art and science of filling the context window with just the right information for the next step.”
IBM’s explainer places two older practices inside it: prompt engineering, which shapes the prompt, and retrieval-augmented generation (RAG), which adds documents to the window. Neither covers the system prompt, message history and tool output together. Context engineering does.
Why does context need engineering at all?
Because more context makes models less reliable, not just more expensive. Chroma’s Context Rot report (July 2025) tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3, and found “performance grows increasingly unreliable as input length grows,” even on simple tasks. Anthropic calls this the model’s “attention budget”: every new token “depletes this budget by some amount.”
Claude Code’s best practices say the same for coding: “Claude’s context window fills up fast, and performance degrades as it fills.”
LangChain’s post borrows four names for what goes wrong, from Drew Breunig:
- Context poisoning: a hallucination makes it into the context.
- Context distraction: the context overwhelms what the model learned in training.
- Context confusion: superfluous context influences the response.
- Context clash: parts of the context disagree.
What goes into an agent’s context?
Six parts, with what Anthropic’s post says good looks like for each:
| Part | What it is | What good looks like |
|---|---|---|
| Instructions | System prompt, CLAUDE.md, AGENTS.md, rules files | The “right altitude”: specific enough to guide behavior, flexible enough to give strong heuristics. The minimal set that fully outlines what you expect |
| Examples | Few-shot examples | A few “diverse, canonical examples”, not a laundry list of edge cases |
| Tools | Tool definitions and their results | A minimal set with little overlap, returning token-efficient results |
| Retrieved knowledge | Files, docs, search results, RAG chunks | Loaded “just in time” through tools, from file paths, queries or links |
| Memory | Notes kept outside the window | Written by the agent as it works, read back after a reset |
| History | Messages, tool calls and tool output | Compacted or cleared. Old raw tool results are the first thing to drop |
On tools, Anthropic’s test is blunt: “If a human engineer can’t definitively say which tool should be used in a given situation, an AI agent can’t be expected to do better.”
What are the main context engineering techniques?
LangChain groups them into four verbs. Anthropic’s post describes the same moves, with Claude Code as the example.
- Write: save context outside the window. Anthropic calls this structured note-taking: “the agent regularly writes notes persisted to memory outside of the context window.” A to-do list or a
NOTES.mdfile is enough. After a reset, the agent reads its notes and carries on. - Select: pull in only what this step needs. Instead of loading everything up front, agents keep “lightweight identifiers (file paths, stored queries, web links, etc.)” and fetch the data with tools when needed. Claude Code is a hybrid: CLAUDE.md files go into context at the start, while glob and grep fetch files on demand.
- Compress: summarize or trim. Compaction summarizes a conversation near the limit and starts a new window with the summary. Anthropic says Claude Code’s summary keeps “architectural decisions, unresolved bugs, and implementation details”, drops redundant tool output, and continues with the five most recently accessed files. Clearing old tool results is “one of the safest lightest touch forms of compaction.”
- Isolate: split work across sub-agents. Each sub-agent has a clean window, may use tens of thousands of tokens, and returns “a condensed, distilled summary of its work (often 1,000-2,000 tokens).” The cost is tokens: Anthropic’s multi-agent research post reports agents using about 4× the tokens of chat, and multi-agent systems about 15×.
Anthropic’s guide to choosing: compaction for long back-and-forth, note-taking for iterative development with clear milestones, sub-agents for research where parallel exploration pays off.
How do I apply context engineering to a coding agent?
The commands below are Claude Code’s, from its best practices, costs and context window pages. The ideas apply to any coding agent.
-
Look before you cut. Run
/contextto see what is using the window. Our Claude Code context window guide goes through what loads at startup. -
Keep always-loaded instructions short. CLAUDE.md loads every session, so put only what applies broadly there (see what to put in CLAUDE.md). Procedures you need sometimes belong in skills: only a one-line description loads at startup, and the full skill loads when used. Of auto memory’s
MEMORY.md, the first 200 lines or 25KB load. -
Trim tool overhead. MCP tool definitions are deferred by default, so only tool names and server instructions enter context until a tool is used. The docs still suggest CLI tools such as
ghwhere they exist, and disabling servers you are not using under/mcp. -
Clear between tasks.
/clearresets the window. If you have corrected Claude more than twice on the same issue, the docs say the context is “cluttered with failed approaches”: clear it and start over with a better prompt. -
Steer compaction. Run
/compact Focus on the API changes, or leave a standing instruction in CLAUDE.md:# Compact instructions When you are using compact, please focus on test output and code changes -
Send investigations to sub-agents. Ask Claude to “use subagents to investigate X”. The file reading happens in a separate window and only a summary comes back.
-
Write down what the next session needs. A compaction summary lives and dies with its session. A decision you will need next week belongs in a file or a page, not only in the chat.
What about context shared across several agents?
Every technique above is about one agent’s window. Most teams run several: Claude Code in one repo, Codex in another, Cursor on a teammate’s laptop, a chat assistant, and sub-agents spun up for an afternoon. Each has its own instruction files, its own memory and its own history. We run several Hermes profiles on one Mac, and each keeps its own memory: what one profile learns, the others never see unless it lands somewhere they all read.
That adds three problems most guides skip:
- Knowledge stuck in one agent. A decision Claude Code worked out sits in its auto memory or a compaction summary. Codex starts from zero.
- Clash between agents. Two agents keep two copies of the same fact. One gets updated, and now they disagree. It is Breunig’s context clash, across agents instead of inside one window.
- Stale context. A note written once and never revised is poisoning on a delay: the agent loads it with full confidence.
The fixes reuse the same ideas, applied to a shared store:
- One source of truth outside every window. Notes go somewhere every agent can read, not into any one agent’s memory. It is Anthropic’s note-taking with more than one reader.
- Select just in time. Agents search for and read the page a task needs instead of loading everything at startup. Instruction files point to the store; they do not copy it.
- One page per topic, edited in place. Update the page that exists rather than appending a new note. That is how clash is avoided.
- Record who wrote what, and why. The agent’s name and a reason on every change, so a wrong note can be traced and fixed.
- Review what may be out of date. Look at pages nobody has touched while the things they describe changed.
If all your agents work in one repository, a well-kept AGENTS.md may be all you need. Beyond that, see agent memory, our roundup of MCP memory servers, and Karpathy’s LLM wiki pattern: linked pages that agents keep up to date.
Share it with your other agents
Dexio is a hosted wiki that AI agents read and write through MCP. Claude Code, Codex, Cursor, Hermes, OpenClaw, Claude, ChatGPT and other MCP clients connect to the same wiki of linked markdown pages, and you see what your agents know as a page graph at app.dexio.wiki. It maps onto the practices above:
- Write and select. Agents use
search_pagesandread_pageto load a page when a task needs it, andwrite_page,edit_pageandappend_pageto record what they learn. Dexio tells connected agents to check the wiki before answering questions about your work, and to record decisions, findings and how things work on the page that covers them. - Who wrote what. Every change names the agent that made it and can carry a note saying why.
page_historyand each page’s History tab show it. - Kept current.
wiki_healthreports pages that may be out of date, and missing pages that many pages link to.
Point every agent’s instruction file at it instead of copying knowledge into each one:
## Shared wiki
- Before you work on a topic, search the Dexio wiki with `search_pages` and read what is there.
- When you work out something that will be needed again, write it down.
Update the page that exists instead of adding a duplicate.
- Put your agent name in the `agent` field on every change.
To connect an agent, send it this:
Connect yourself to my Dexio wiki. The steps are at https://dexio.wiki/agents.md: read the whole file, not a summary, and follow them.
Dexio is free for one person; Team is $10 and Business $20 a member a month, and agents never count as members. It is open source under the AGPL, and any wiki downloads as markdown. Setup per agent is in the guides, starting with Claude Code MCP. To see an existing folder of markdown as a graph first, try the wiki graph viewer.
FAQ
Is context engineering just prompt engineering renamed? The Prompting Guide says prompt engineering “is now being rebranded as context engineering.” The difference for agents: a prompt is written once, while an agent in a loop “generates more and more data that could be relevant for the next turn of inference,” so its context is curated again every turn.
Do million-token context windows make it unnecessary? No. Chroma’s tests show reliability falling as input grows, and Anthropic expects that “context windows of all sizes will be subject to context pollution and information relevance concerns.”
What is context rot? Anthropic’s definition: “as the number of tokens in the context window increases, the model’s ability to accurately recall information from that context decreases.”
Sources
- https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- https://www.anthropic.com/engineering/multi-agent-research-system
- https://www.langchain.com/blog/context-engineering-for-agents
- https://www.ibm.com/think/topics/context-engineering
- https://www.promptingguide.ai/guides/context-engineering-guide
- https://research.trychroma.com/context-rot
- https://code.claude.com/docs/en/best-practices
- https://code.claude.com/docs/en/costs
- https://code.claude.com/docs/en/context-window
- https://dexio.wiki/docs/
- https://dexio.wiki/pricing/
- https://dexio.wiki/agents.md