Agent memory: how AI agents remember, and which kind to use
Updated
Agent memory is whatever lets an AI agent use in one session what it learned in another. The model does not remember on its own: each session starts with a fresh context window, so anything the agent should know has to be saved somewhere and loaded back in. The main kinds are memory files, the chat apps’ built-in memory, memory APIs, retrieval over documents (RAG) and a wiki the agent maintains, and most setups use more than one.
Why do AI agents need memory?
A language model works from its context window: the instructions, files and conversation in front of it now. When the session ends, that context is gone. Anthropic’s docs say it directly: “Each Claude Code session begins with a fresh context window.” The OpenClaw docs say the model only remembers what gets saved to disk.
So memory is two jobs: writing down what should last, and putting the right part of it back into context at the right time. The kinds below differ in who writes, what gets stored and how it comes back.
“LLM memory” can also mean what a model learned in training. The paper that introduced RAG calls this parametric memory and notes that updating it is still an open problem. You cannot add your own work to it, so this guide covers memory kept outside the model.
Memory files the agent loads every session
The simplest kind is a markdown file that loads into context at the start of every session.
- CLAUDE.md holds instructions you write for Claude Code. Claude Code also keeps auto memory, notes Claude writes for itself; the first 200 lines or 25KB of its index load each session, on one machine. See the CLAUDE.md guide and Claude Code memory.
- AGENTS.md is an open format for the same job, read by Codex, Cursor, Claude Code and other coding agents. See the AGENTS.md guide.
- Hermes Agent keeps
MEMORY.mdfor its own notes (2,200 characters) andUSER.mdfor your preferences (1,375). It edits both with its memory tool, and a search over past sessions covers the rest. See Hermes Agent memory. - OpenClaw loads
MEMORY.md, its durable facts and decisions, at the start of a session, and searches its daily notes withmemory_search. See OpenClaw memory.
Good at: rules and preferences the agent needs every time. You can read and edit the files, and project files go into git with the code.
Limits: they load in full, so they must stay short. Anthropic recommends under 200 lines per CLAUDE.md, because longer files use more context and reduce adherence. Most belong to one agent on one machine.
Built-in memory in Claude and ChatGPT
Claude builds memory from your chats. It saves things on its own, and you can tell it to “remember this”. Each project has its own memory, and you can view and edit what it keeps. On Team and Enterprise plans an owner has to turn memory on. See Claude memory.
ChatGPT has saved memories, details you ask it to keep or that it saves as useful, and a setting to reference chat history. OpenAI notes that what it takes from chat history can change as it updates what is most useful to remember. See ChatGPT memory.
Good at: personal context with no setup: your role, your projects, how you like answers.
Limits: it lives in one app and your account there. A coding agent in your terminal does not see it, and it holds context about you rather than documents about your work.
Memory APIs: Mem0, Zep and Letta
These are services for developers building agents. Your code sends them conversations or data and asks for relevant memory when the agent needs it.
- Mem0 runs your messages through an LLM that pulls out key facts, decisions or preferences, scoped by user, agent, app or run, and returns the most relevant ones for a query. It is a managed platform or open source.
- Zep builds temporal Context Graphs of facts, relationships and changes over time from chat messages and business data. When new data invalidates a fact, it records when.
- Letta has the agent manage its own memory as markdown files in a git repository. Files under
system/are in the prompt every turn; the agent reads the rest as needed.
Good at: memory inside a product you build, such as remembering each of your users.
Limits: you build it into your own application. Mem0 and Zep store short facts or graph edges, which a person reads through an API or dashboard rather than as documents.
RAG: retrieval over your documents
Retrieval-augmented generation indexes your documents as they are, and at question time pulls the relevant chunks into context. The 2020 paper that named it paired a model with a dense vector index of Wikipedia.
Good at: large or fast-changing collections, lookups and answers that point back to a source.
Limits: nothing builds up. As Andrej Karpathy puts it, the model is “rediscovering knowledge from scratch on every question.” A question that spans five documents means piecing the fragments together every time.
An LLM wiki the agent maintains
In Karpathy’s LLM wiki pattern, the agent builds and maintains a wiki of interlinked markdown pages between you and your raw sources. When you add a source, the agent works it into the existing pages: it updates entity pages, revises summaries and notes where new data contradicts old claims. Good answers get filed back as new pages. In his words: “You read it; the LLM writes it.”
Good at: knowledge that accumulates, such as research, decisions and how systems fit together, in pages a person can read and follow by link.
Limits: the agent writes its mistakes into pages as readily as its findings, so read what it writes. Upkeep costs tokens, and a wiki in a local folder belongs to one machine.
Agent memory types compared
| Kind | Who writes it | What it holds | Can a person read it | Shared across agents and machines |
|---|---|---|---|---|
| Memory files | You, or the agent | Instructions, preferences, short notes | Yes | Project files through git; the rest stays on one machine |
| Chat app memory | The app, from your chats | Context about you | Yes, in settings | No, one app and one account |
| Memory APIs | The service’s model, or the agent | Facts, graph edges or memory files | Through an API or dashboard | Across whatever your app connects |
| RAG | Nobody; documents are indexed | Chunks of your sources | The sources, not the index | Anything that queries the index |
| LLM wiki | The agent, guided by you | Linked pages: summaries, decisions, research | Yes | One machine as a folder; every agent when hosted |
Which kind of agent memory should you use?
Pick by the question each one answers:
- What must the agent do every time? A memory file.
- Who am I, and what am I working on? The chat app’s memory.
- What does each user of my product want? A memory API.
- What does this large set of documents say? RAG.
- What have we worked out so far? A wiki.
Many setups combine them. A coding agent keeps its rules in CLAUDE.md or AGENTS.md and points to a wiki for everything longer. Karpathy’s gist works the same way: a schema file, such as CLAUDE.md or AGENTS.md, tells the agent how to keep the wiki.
Where a shared wiki fits
Three cases push you past files and app memory:
- Several agents and machines. A Claude chat, Claude Code on your laptop and Hermes on a server each remember separately.
- Knowledge that is more than facts. Why you chose a vendor, what an investigation found, how three services fit together. That takes pages, not lines.
- People who need to read it. To check what your agents believe, you need it in words you can open.
A shared wiki with Dexio
Dexio is a hosted wiki for AI agents. Your agents list, read, search, write, edit, move and link markdown pages over MCP, and you see the link graph and broken links at https://app.dexio.wiki. Every agent you connect works on the same wiki, and every change is kept with the agent that made it. It is free for one person.
Hermes, OpenClaw, Claude Code and other agents that run commands set themselves up. Send your agent this message:
Connect yourself to my Dexio wiki. Read the steps with curl -s https://dexio.wiki/agents.md and follow them.
Claude and ChatGPT connect as apps instead. The steps for each agent are in the guides.
Sources
- https://code.claude.com/docs/en/memory
- https://agents.md
- https://hermes-agent.nousresearch.com/docs/user-guide/features/memory
- https://docs.openclaw.ai/concepts/memory
- https://support.claude.com/en/articles/11817273-use-claude-s-chat-search-and-memory-to-build-on-previous-context
- https://help.openai.com/en/articles/8590148-memory-in-chatgpt
- https://docs.mem0.ai/overview
- https://docs.mem0.ai/core-concepts/memory-operations/add
- https://help.getzep.com/concepts
- https://docs.letta.com/agent-sdk/memory
- https://docs.letta.com/concepts/memfs
- https://arxiv.org/abs/2005.11401
- https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
- https://dexio.wiki/agents.md
- https://dexio.wiki/docs