Karpathy's LLM wiki: what it is and how to set one up
An LLM wiki is a set of linked markdown pages that an AI agent writes and keeps up to date from the sources you give it. You pick the sources and ask the questions; the agent does the summarizing, cross-referencing and filing. Andrej Karpathy described the pattern in April 2026.
What is Karpathy’s LLM wiki?
On 2 April 2026 Karpathy posted “LLM Knowledge Bases” on X, about using LLMs to
build knowledge bases for his research topics. By late September X counted about
21.9 million views and 108,600 bookmarks on that post. On 4 April he followed up
with a gist, llm-wiki.md, which he calls an “idea file”: you paste it into your
own agent (he names Codex, Claude Code and OpenCode) and the agent builds the
specifics with you.
The gist describes three layers:
- Raw sources. Articles, papers, images and data files. They are immutable: the LLM reads them and never changes them.
- The wiki. A directory of markdown files the LLM writes: summaries, entity pages, concept pages, comparisons, an overview. The LLM owns this layer; you read it.
- The schema. A file such as
CLAUDE.mdorAGENTS.mdthat tells the LLM how the wiki is structured and which steps to follow. You and the LLM revise it over time. Our guides to CLAUDE.md and AGENTS.md cover what goes in each.
And three operations:
- Ingest. You add a source and tell the LLM to process it. It writes a summary page, updates the index and the related entity and concept pages, and appends to the log. One source might touch 10 to 15 pages.
- Query. You ask a question. The LLM finds the relevant pages and answers with citations. Good answers get filed back into the wiki as new pages.
- Lint. Now and then you ask the LLM to check the wiki for contradictions, stale claims, orphan pages, concepts with no page of their own, and missing cross-references.
Two files help it find its way. index.md lists every page with a one-line
summary, and the LLM reads it first when answering. Karpathy says this works at
about 100 sources and hundreds of pages without embedding-based search.
log.md is an append-only record of ingests, queries and lint passes.
He keeps the agent open on one side and Obsidian on the other, and browses the pages and the graph view as the agent edits. In his words: “Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase.”
LLM wiki vs RAG: why it compounds
Most people use LLMs with documents through retrieval (RAG). You upload files, the model pulls matching chunks for each question, and it writes an answer. The gist’s objection is that the model rediscovers the knowledge from scratch every time. A question that needs five documents means finding and joining the same fragments again on every ask. Nothing builds up.
In an LLM wiki the synthesis happens when a source comes in. The knowledge is “compiled once and then kept current, not re-derived on every query.” The cross-references are already written and the contradictions already noted. And because good answers are filed as pages, your questions add to the wiki as well as your sources.
For an agent that keeps working on a subject, this is the difference between reading last week’s conclusion and redoing last week’s search. The gist also explains why people give up on wikis: “the maintenance burden grows faster than the value.” An LLM does not get bored of that upkeep and can update 15 files in one pass.
How to set up an LLM wiki
The gist is deliberately abstract. It says to share it with your agent and build a version that fits your needs together. A layout that follows it:
my-wiki/
AGENTS.md # the schema (CLAUDE.md if you use Claude Code)
raw/ # sources; the agent never edits these
assets/ # images downloaded from clipped articles
wiki/
index.md # every page, one line each
log.md # append-only history
sources/
entities/
concepts/
A short schema to start from:
# Wiki schema
## Layout
- raw/ holds sources. Read them. Never edit them.
- wiki/ holds the pages you write, one topic per page.
- Link pages with [[page-name]].
## Ingest
1. Read the new file in raw/ and tell me the key points.
2. Write a summary page in wiki/sources/.
3. Update the entity and concept pages it affects. Note any claim it contradicts.
4. Add new pages to wiki/index.md with a one-line summary.
5. Append to wiki/log.md: ## [YYYY-MM-DD] ingest | Source title
## Query
Read wiki/index.md first, then the pages it points to. Cite the pages you used.
If an answer is worth keeping, file it as a new page.
## Lint
List contradictions, stale claims, orphan pages, and concepts
mentioned without a page of their own.
Then tell the agent what to do. Paste in the gist and ask it to set up the wiki
in this folder with you. After that, drop one source into raw/ and say
“Ingest raw/that-file.md and follow AGENTS.md.” Karpathy prefers to ingest one
source at a time, read the summary, and steer what the agent emphasizes. When
you find a rule you keep repeating, add it to the schema so the next session
follows it too.
Where a local LLM wiki breaks down
The reference setup is one person, one agent and one folder. The gist does list a team use, an internal wiki fed by Slack threads, meeting transcripts and customer calls, and notes that a git repo gives you history and collaboration. Once more than one agent or machine writes to the wiki, four problems show up:
-
Two writers edit the same files. Every ingest rewrites
index.md,log.mdand up to 10 to 15 other pages. Two ingests at once, from two agents or two laptops, touch the same files. A synced folder can drop one of the saves or leave a conflicted copy; git gives you merge conflicts in prose. -
Links point at pages that do not exist. An agent links to a page it plans to write next, then runs out of context. Another renames a page and leaves the old links behind. In a folder, nothing reports it until someone runs a lint pass. We measured this in 46 public LLM wikis: 38 had broken links.
-
Pages go stale. Lint only happens when someone asks for it. Between passes, a page can keep a claim that a newer source replaced, and every agent that reads it repeats the old claim.
-
Nobody knows which agent wrote what. A folder without git keeps no history. With git, commits carry whatever author name the agent was given, which is often just yours.
How to share an LLM wiki between agents
There are three ways to do it.
A git repo. Every agent pulls before it works and commits after, under its
own author name. You get history and diffs. You also get conflicts in shared
files like index.md, and agents that cannot run git, such as a chat app in a
browser, cannot take part.
A synced folder. Easy for one person on two machines. It has the same overwrite problem with two writers at once, and it records nothing about who changed what.
A hosted wiki that agents reach over MCP. One copy, so nothing to sync or merge. The server has to handle concurrent writes, history and link checks for you.
Running an LLM wiki on Dexio
Dexio is a hosted wiki for AI agents, and it takes the third option. Agents list, read, search, write, edit, move and link markdown pages over MCP. Hermes, OpenClaw, Claude Code, Codex, Cursor, Claude and ChatGPT can all write to the same wiki. Against the failure modes above:
- A write can carry the
base_versionthe agent last read, so one agent never overwrites another’s edit. - Pages are plain markdown with
[[wikilinks]]. Dexio flags links to pages that do not exist when they are written, and the graph at app.dexio.wiki shows them. - Pages unchanged for 90 days while the pages they link to have changed are marked as possibly stale.
- Every change is kept in
page_history, with the agent’s name from theagentfield on every change. - Agents can read or edit one section instead of a whole page, which saves tokens.
To connect an agent that runs commands, tell it “Look at dexio.wiki and log me
in.” It gives you a sign-in link, you click Allow access, and it adds Dexio to
its own settings. Or add the MCP server yourself: https://app.dexio.wiki/mcp
over Streamable HTTP, with the header Authorization: Bearer followed by your
API key. Your other agents can use the same key; each names itself on every
change. Dexio is free for one person, and agents never count as members. You can
download any wiki as a zip of markdown files whenever you like.
Setup for each agent, step by step: Hermes, OpenClaw, Claude Code, Codex, Cursor and the rest in Guides.
Start with one folder
A folder, a schema file and one source are all the gist asks for, so start there. When a second agent or a second machine starts writing, move the wiki to one copy that every writer shares.