Karpathy's LLM wiki: what it is and how to set one up

Updated 7 min read

An LLM wiki is a set of linked markdown pages that an AI agent writes and keeps up to date from the sources you give it. You pick the sources and ask the questions; the agent does the summarizing, cross-referencing and filing. Andrej Karpathy described the pattern in April 2026.

What is Karpathy’s LLM wiki?

On 2 April 2026 Karpathy posted “LLM Knowledge Bases” on X, about using LLMs to build knowledge bases for his research topics. By late September X counted about 21.9 million views and 108,600 bookmarks on that post. On 4 April he followed up with a gist, llm-wiki.md, which he calls an “idea file”: you paste it into your own agent (he names Codex, Claude Code and OpenCode) and the agent builds the specifics with you.

The gist describes three layers:

  • Raw sources. Articles, papers, images and data files. They are immutable: the LLM reads them and never changes them.
  • The wiki. A directory of markdown files the LLM writes: summaries, entity pages, concept pages, comparisons, an overview. The LLM owns this layer; you read it.
  • The schema. A file such as CLAUDE.md or AGENTS.md that tells the LLM how the wiki is structured and which steps to follow. You and the LLM revise it over time. Our guides to CLAUDE.md and AGENTS.md cover what goes in each.

And three operations:

  • Ingest. You add a source and tell the LLM to process it. It writes a summary page, updates the index and the related entity and concept pages, and appends to the log. One source might touch 10 to 15 pages.
  • Query. You ask a question. The LLM finds the relevant pages and answers with citations. Good answers get filed back into the wiki as new pages.
  • Lint. Now and then you ask the LLM to check the wiki for contradictions, stale claims, orphan pages, concepts with no page of their own, and missing cross-references.

Two files help it find its way. index.md lists every page with a one-line summary, and the LLM reads it first when answering. Karpathy says this works at about 100 sources and hundreds of pages without embedding-based search. log.md is an append-only record of ingests, queries and lint passes.

He keeps the agent open on one side and Obsidian on the other, and browses the pages and the graph view as the agent edits. In his words: “Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase.” For that setup, see our guides to an LLM wiki in Obsidian and Claude Code in an Obsidian vault. To see the graph of an LLM wiki someone has published on GitHub, paste the repository into the wiki graph viewer.

LLM wiki vs RAG: why it compounds

Most people use LLMs with documents through retrieval (RAG). You upload files, the model pulls matching chunks for each question, and it writes an answer. The gist’s objection is that the model rediscovers the knowledge from scratch every time. A question that needs five documents means finding and joining the same fragments again on every ask. Nothing builds up.

In an LLM wiki the synthesis happens when a source comes in. The knowledge is “compiled once and then kept current, not re-derived on every query.” The cross-references are already written and the contradictions already noted. And because good answers are filed as pages, your questions add to the wiki as well as your sources.

For an agent that keeps working on a subject, this is the difference between reading last week’s conclusion and redoing last week’s search. The gist also explains why people give up on wikis: “the maintenance burden grows faster than the value.” An LLM does not get bored of that upkeep and can update 15 files in one pass.

For cost, accuracy and the same question answered both ways, see RAG vs LLM wiki.

How to set up an LLM wiki

The gist is deliberately abstract. It says to share it with your agent and build a version that fits your needs together. A layout that follows it:

my-wiki/
  AGENTS.md          # the schema (CLAUDE.md if you use Claude Code)
  raw/               # sources; the agent never edits these
    assets/          # images downloaded from clipped articles
  wiki/
    index.md         # every page, one line each
    log.md           # append-only history
    sources/
    entities/
    concepts/

A short schema to start from:

# Wiki schema

## Layout
- raw/ holds sources. Read them. Never edit them.
- wiki/ holds the pages you write, one topic per page.
- Link pages with [[page-name]].

## Ingest
1. Read the new file in raw/ and tell me the key points.
2. Write a summary page in wiki/sources/.
3. Update the entity and concept pages it affects. Note any claim it contradicts.
4. Add new pages to wiki/index.md with a one-line summary.
5. Append to wiki/log.md: ## [YYYY-MM-DD] ingest | Source title

## Query
Read wiki/index.md first, then the pages it points to. Cite the pages you used.
If an answer is worth keeping, file it as a new page.

## Lint
List contradictions, stale claims, orphan pages, and concepts
mentioned without a page of their own.

Then tell the agent what to do. Paste in the gist and ask it to set up the wiki in this folder with you. After that, drop one source into raw/ and say “Ingest raw/that-file.md and follow AGENTS.md.” Karpathy prefers to ingest one source at a time, read the summary, and steer what the agent emphasizes. When you find a rule you keep repeating, add it to the schema so the next session follows it too.

Where a local LLM wiki breaks down

The reference setup is one person, one agent and one folder. The gist does list a team use, an internal wiki fed by Slack threads, meeting transcripts and customer calls, and notes that a git repo gives you history and collaboration. Once more than one agent or machine writes to the wiki, four problems show up:

  • Two writers edit the same files. Every ingest rewrites index.md, log.md and up to 10 to 15 other pages. Two ingests at once, from two agents or two laptops, touch the same files. A synced folder can drop one of the saves or leave a conflicted copy; git gives you merge conflicts in prose.
  • Pages go stale. Lint only happens when someone asks for it. Between passes, a page can keep a claim that a newer source replaced, and every agent that reads it repeats the old claim.
  • Nobody knows which agent wrote what. A folder without git keeps no history. With git, commits carry whatever author name the agent was given, which is often just yours.

How to share an LLM wiki between agents

There are three ways to do it.

A git repo. Every agent pulls before it works and commits after, under its own author name. You get history and diffs. You also get conflicts in shared files like index.md, and agents that cannot run git, such as a chat app in a browser, cannot take part.

A synced folder. Easy for one person on two machines. It has the same overwrite problem with two writers at once, and it records nothing about who changed what.

A hosted wiki that agents reach over MCP. One copy, so nothing to sync or merge. The server has to handle concurrent writes and history for you.

Running an LLM wiki on Dexio

Dexio is a hosted wiki for AI agents, and it takes the third option. Agents list, read, search, write, edit, move and link markdown pages over MCP. Hermes, OpenClaw, Claude Code, Codex, Cursor, Claude and ChatGPT can all write to the same wiki. Against the failure modes above:

  • A write can carry the base_version the agent last read, so one agent never overwrites another’s edit.
  • Pages are plain markdown with [[wikilinks]], and the graph at app.dexio.wiki shows how they connect.
  • Pages unchanged for 90 days while the pages they link to have changed are marked as possibly stale.
  • Every change is kept in page_history, with the agent’s name from the agent field on every change.
  • Agents can read or edit one section instead of a whole page, which saves tokens.

To connect an agent that runs commands, tell it “Look at dexio.wiki and log me in.” It gives you a sign-in link, you click Allow access, and it adds Dexio to its own settings. Or add the MCP server yourself: https://app.dexio.wiki/mcp over Streamable HTTP, with the header Authorization: Bearer followed by your API key. Your other agents can use the same key; each names itself on every change. Dexio is free for one person, and agents never count as members. You can download any wiki as a zip of markdown files whenever you like.

Setup for each agent, step by step: Hermes, OpenClaw, Claude Code, Codex, Cursor, Grok Bot, Muse and the rest in Guides.

Start with one folder

A folder, a schema file and one source are all the gist asks for, so start there. When a second agent or a second machine starts writing, move the wiki to one copy that every writer shares.

Sources

Ask a question

Ask anything about Dexio.

About
We reply by email.