RAG vs LLM wiki: cost, accuracy and when to use each
Updated 10 min read
RAG and an LLM wiki do their work at different times. RAG indexes your documents as they are and assembles an answer from retrieved chunks every time you ask; an LLM wiki has the model do the synthesis when a source comes in, so a question reads conclusions that already exist. In the one public head-to-head with numbers, RAG was cheaper, faster and slightly more accurate on lookups, while a wiki is built for knowledge that accumulates and that people need to read. Most setups that last end up using both.
What is the difference between RAG and an LLM wiki?
RAG (retrieval-augmented generation) splits your documents into chunks, indexes them (with embeddings, keywords or both), and at question time pulls the best-matching chunks into the model’s context. Nothing is written back. Andrej Karpathy’s gist describes it as the model “rediscovering knowledge from scratch on every question” and puts NotebookLM, ChatGPT file uploads and most RAG systems in this group.
An LLM wiki, in the gist’s version, has the model read each new source,
write a summary page, update the 10 to 15 entity and concept pages it touches,
note contradictions, and add the page to index.md. At question time it reads
the index, opens the relevant pages and answers with citations. Good answers
are filed back as new pages. Our post on Karpathy’s LLM wiki
covers the pattern and setup.
| RAG | LLM wiki | |
|---|---|---|
| When the model does the synthesis | On every question | When a source arrives, and on lint passes |
| What is stored | Chunks of raw documents, plus an index | Pages the model wrote, plus the raw sources |
| How a question finds knowledge | Similarity or keyword search over chunks | An index file, links, or search over pages |
| What a person can read | The sources | The sources and the pages |
| Where mistakes live | In one answer | In a page, until someone fixes it |
| What grows with use | The index | The pages, including filed answers |
The same question answered both ways
An illustration, not a benchmark. An agent helps a team with its billing service, and has three sources:
raw/2026-03-design-review.md: “We chose Postgres for billing: one database, strong transactions.”raw/2026-06-outage-notes.md: “Ledger writes move to a separate store after the 12 June outage. Invoices stay in Postgres.”raw/2026-08-standup.md: “Ledger migration done on 14 August.”
Someone asks: “What database does billing use now, and why?”
With RAG. The question is matched against the chunks. The March chunk scores best: it has “billing”, “Postgres” and “chose”. The June chunk talks about “ledger” and “outage” and may rank lower. The August line mentions neither billing nor a database and may not come back at all. The model sees something like:
[1] 2026-03-design-review.md: "We chose Postgres for billing: one database, strong transactions..."
[2] 2026-06-outage-notes.md: "Ledger writes move to a separate store after the 12 June outage..."
A careful model answers “Postgres, with the ledger moving out after June” and cannot say whether the move happened. A careless one answers “Postgres”. The next person to ask gets the same search and the same assembly.
With an LLM wiki. The work happened at ingest. When the June notes came in,
the agent edited services/billing.md; when the standup came in, it edited the
page again and appended to log.md. The page now reads:
# Billing service
## Storage
- Invoices: Postgres (chosen 2026-03, see [[decisions/billing-database]]).
- Ledger: separate store since 2026-08-14. Moved after the
[[incidents/2026-06-billing-outage]]; migration done per the 2026-08 standup.
The agent reads index.md, opens this page, and answers with all three facts
and a link to each source. Nothing depends on the August line sharing words with
the question.
The flip side: if the agent had misread the June notes and written “billing moved off Postgres”, every answer after that would repeat it, and no retrieval step would surface the original wording. That is the trade in one example. RAG redoes the synthesis each time and can miss pieces; a wiki does it once and keeps its mistakes as faithfully as its findings.
Which is cheaper, RAG or an LLM wiki?
RAG costs little to write (chunking and embedding) and a modest amount per question. A wiki spends model tokens on every ingest, and per question it costs whatever page text the model reads.
The only public head-to-head of the two with published numbers that we found, as
of October 2026, is a practitioner’s evaluation
published in May 2026, in Chinese, on the osisdie blog. Their method: one
internal knowledge base, a hybrid RAG pipeline (BM25 plus dense vectors, merged
with reciprocal rank fusion) against a wiki pipeline in which a model reads
index.md, picks up to four pages and answers from them. The wiki had 60 pages
built from the same data; the vector store held 1,402 points. They graded 13
questions by hand, a single-tenant subset of 43. Their numbers:
| Measure (their run) | Hybrid RAG | LLM wiki |
|---|---|---|
| Correct answers | 13 of 13 | 12 of 13 |
| Tokens per question, average | 1,407 | 13,055 |
| Latency, median | 3.3 s | 8.2 s |
| Latency, 95th percentile | 24.3 s | 27.3 s |
| LLM calls per question | 2 | 4 |
| Building the wiki | None | About 194,000 tokens, once |
Adding one new source took 3,000 to 5,000 tokens in their setup. The two extra calls are by design: the wiki pipeline replaces search with two model steps, pick pages and then answer. On the one out-of-scope question the wiki was faster (1.6 s against 2.3 s), because the index step found nothing and stopped.
Two limits they state themselves: 13 questions is a small set, and they were mostly direct FAQ lookups, where RAG is strong. The case a wiki is built for, a question that needs several documents put together, was not in it.
Which is more accurate?
Neither by default. The wiki’s one miss in that evaluation is worth reading. Asked whether a policy could be transferred, it said yes and listed another tenant’s quotas, with a self-reported confidence of 0.9. The right answer for this tenant was no. The cause was at write time: pages were grouped by entity name rather than by tenant and entity, so two tenants’ policies were merged into one page. They found two more patterns:
- A small model writing pages added noise (a stray line of Arabic and English in a Chinese answer). It was saved in the page, so every later answer repeated it.
- Short paraphrased questions got one-line answers that dropped exceptions the RAG answer listed.
The wiki also reported higher average confidence (0.862 against 0.808) while scoring lower. Their conclusion: hybrid RAG stays the default, the wiki becomes an opt-in mode per request, and before switching, add a reranker, contextual retrieval, intent routing and tenant filters on the RAG side.
Two other results point the same way from different angles. Letta put LoCoMo conversation histories in files and gave a gpt-4o-mini agent file tools, grep and semantic search; it scored 74.0%, against the 68.5% Mem0 reported for its best graph variant. Letta’s reading is that how the agent searches (rewriting queries, searching again) matters more than the storage behind it. And a study of AGENTS.md files (arXiv 2602.11988) found that context files did not generally raise coding agents’ task success and raised inference cost by more than 20%. Repository overviews did not help; specific instructions were followed. Pre-written summaries are not free accuracy. What to load into context, and when, is the subject of our context engineering guide.
Which stays fresher?
RAG is current as soon as a document is indexed, and nobody rewrites anything. But it has no notion of “replaced”: the March decision and the June change sit in the index with equal standing, and the model sorts them out at answer time, every time.
A wiki is only as current as its last ingest. If sources arrive and nobody tells the agent to process them, the pages lag. In exchange, a contradiction is handled once, when the new source arrives, and the page says what changed and when. The gist’s lint step, which looks for contradictions, stale claims and orphan pages, only runs when someone asks for it.
Where does each stop scaling?
Karpathy says the index file “works surprisingly well at moderate scale (~100 sources, ~hundreds of pages)” without embedding-based retrieval. Past that he suggests proper search over the pages and names qmd, a local search engine for markdown with hybrid BM25 and vector search, LLM re-ranking, a CLI and an MCP server. In other words, a large wiki needs retrieval again; it retrieves pages instead of raw chunks.
RAG scales to collections nobody would want summarized: millions of chunks, documents that change hourly, archives searched once a year. Its limit runs the other way. Questions that need many documents at once get top-k fragments.
Boundaries are a wiki-specific risk. A model writing pages merges whatever looks alike unless the schema says otherwise. If your data has partitions (customers, tenants, teams that must not see each other’s information), put the partition in the page path or frontmatter and filter on it, or keep separate wikis.
When should you use RAG and when an LLM wiki?
Use RAG when:
- the collection is large, or changes faster than anyone would re-ingest it;
- questions are lookups with the answer in one place: FAQs, policies, docs;
- answers must trace to the original wording, or be filtered per customer;
- cost and latency per question matter (about a ninth of the wiki’s tokens in the evaluation above).
Use an LLM wiki when:
- knowledge builds up over weeks: research, decisions, how systems fit together;
- questions span several sources and similar ones come up again;
- people need to read, and correct, what the agent believes;
- several agents should start from the same conclusions instead of re-deriving them.
Our agent memory guide sets both next to memory files and memory APIs.
Can you combine RAG and an LLM wiki?
Yes, and most working setups do. The gist itself adds search once the wiki grows. Four patterns:
- Search over the pages. Keep the wiki, and index its pages for keyword or vector search instead of relying on one index file.
- Wiki first, raw sources second. Answer from the pages; when they do not
cover the question, fall back to RAG over
raw/and file the answer as a new page, so the next person reads it. - Route by question. The evaluation above sends lookups to hybrid RAG and keeps the wiki as a per-request mode.
- Lint on a schedule. Merged partitions and noise saved in pages are what a lint pass is for, if it runs.
Several tools in our LLM wiki GitHub roundup already pair a wiki with vector search. For memory that agents reach as an MCP server, see the best MCP memory servers.
Share the wiki with your other agents
A wiki pays off when the next session, and the next agent, starts from what was already worked out. That only happens if every agent reads the same pages.
Dexio is a hosted wiki that agents read and write over
MCP. Claude Code, Codex, Cursor, Hermes Agent, OpenClaw, Claude and ChatGPT
connect to one wiki of linked markdown pages, and you see what your agents know
as a page graph. Raw files such as PDFs, decks and spreadsheets can sit beside
the pages, and search_pages searches their text too. Every change records
which agent made it. Dexio is open source (AGPL-3.0), free for one person, and
$10 (Team) or $20 (Business) a member a month.
Dexio’s search matches every word of a question, best matches first. For semantic search across a large document collection, a RAG stack is the better pick; run it beside the wiki and have agents file what they conclude. To connect an agent that runs commands, tell it “Look at dexio.wiki and log me in.” Setup per agent: Claude Code, Hermes Agent and the rest in Guides.
FAQ
Does an LLM wiki replace RAG? Not on the evidence published so far. It replaces the synthesis RAG redoes on every question, not retrieval: once a wiki outgrows its index file, finding the right pages is a retrieval problem again.
Does an LLM wiki need embeddings?
No. Karpathy’s version reads index.md and follows links. Larger wikis add
search, sometimes with vectors.
Is NotebookLM RAG? The gist puts NotebookLM, ChatGPT file uploads and “most RAG systems” in the group that retrieves chunks at question time.
How many sources can an LLM wiki handle? About 100 sources and a few hundred pages with only an index file, per the gist. Beyond that, add search over the pages.
Sources
- https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
- https://osisdie.github.io/blog/2026/hybrid-vs-llm-wiki-eval/
- https://www.letta.com/blog/benchmarking-ai-agent-memory
- https://arxiv.org/abs/2602.11988
- https://affine.pro/blog/what-is-llm-wiki
- https://dev.to/hjarni/karpathys-llm-wiki-is-right-i-just-didnt-want-to-run-it-locally-170m
- https://dexio.wiki/agents.md
- https://dexio.wiki/docs/
- https://dexio.wiki/pricing/