RAG vs LLM wiki: cost, accuracy and when to use each

Updated 10 min read

RAG and an LLM wiki do their work at different times. RAG indexes your documents as they are and assembles an answer from retrieved chunks every time you ask; an LLM wiki has the model do the synthesis when a source comes in, so a question reads conclusions that already exist. In the one public head-to-head with numbers, RAG was cheaper, faster and slightly more accurate on lookups, while a wiki is built for knowledge that accumulates and that people need to read. Most setups that last end up using both.

What is the difference between RAG and an LLM wiki?

RAG (retrieval-augmented generation) splits your documents into chunks, indexes them (with embeddings, keywords or both), and at question time pulls the best-matching chunks into the model’s context. Nothing is written back. Andrej Karpathy’s gist describes it as the model “rediscovering knowledge from scratch on every question” and puts NotebookLM, ChatGPT file uploads and most RAG systems in this group.

An LLM wiki, in the gist’s version, has the model read each new source, write a summary page, update the 10 to 15 entity and concept pages it touches, note contradictions, and add the page to index.md. At question time it reads the index, opens the relevant pages and answers with citations. Good answers are filed back as new pages. Our post on Karpathy’s LLM wiki covers the pattern and setup.

RAG LLM wiki
When the model does the synthesis On every question When a source arrives, and on lint passes
What is stored Chunks of raw documents, plus an index Pages the model wrote, plus the raw sources
How a question finds knowledge Similarity or keyword search over chunks An index file, links, or search over pages
What a person can read The sources The sources and the pages
Where mistakes live In one answer In a page, until someone fixes it
What grows with use The index The pages, including filed answers

The same question answered both ways

An illustration, not a benchmark. An agent helps a team with its billing service, and has three sources:

  • raw/2026-03-design-review.md: “We chose Postgres for billing: one database, strong transactions.”
  • raw/2026-06-outage-notes.md: “Ledger writes move to a separate store after the 12 June outage. Invoices stay in Postgres.”
  • raw/2026-08-standup.md: “Ledger migration done on 14 August.”

Someone asks: “What database does billing use now, and why?”

With RAG. The question is matched against the chunks. The March chunk scores best: it has “billing”, “Postgres” and “chose”. The June chunk talks about “ledger” and “outage” and may rank lower. The August line mentions neither billing nor a database and may not come back at all. The model sees something like:

[1] 2026-03-design-review.md: "We chose Postgres for billing: one database, strong transactions..."
[2] 2026-06-outage-notes.md: "Ledger writes move to a separate store after the 12 June outage..."

A careful model answers “Postgres, with the ledger moving out after June” and cannot say whether the move happened. A careless one answers “Postgres”. The next person to ask gets the same search and the same assembly.

With an LLM wiki. The work happened at ingest. When the June notes came in, the agent edited services/billing.md; when the standup came in, it edited the page again and appended to log.md. The page now reads:

# Billing service

## Storage
- Invoices: Postgres (chosen 2026-03, see [[decisions/billing-database]]).
- Ledger: separate store since 2026-08-14. Moved after the
  [[incidents/2026-06-billing-outage]]; migration done per the 2026-08 standup.

The agent reads index.md, opens this page, and answers with all three facts and a link to each source. Nothing depends on the August line sharing words with the question.

The flip side: if the agent had misread the June notes and written “billing moved off Postgres”, every answer after that would repeat it, and no retrieval step would surface the original wording. That is the trade in one example. RAG redoes the synthesis each time and can miss pieces; a wiki does it once and keeps its mistakes as faithfully as its findings.

Which is cheaper, RAG or an LLM wiki?

RAG costs little to write (chunking and embedding) and a modest amount per question. A wiki spends model tokens on every ingest, and per question it costs whatever page text the model reads.

The only public head-to-head of the two with published numbers that we found, as of October 2026, is a practitioner’s evaluation published in May 2026, in Chinese, on the osisdie blog. Their method: one internal knowledge base, a hybrid RAG pipeline (BM25 plus dense vectors, merged with reciprocal rank fusion) against a wiki pipeline in which a model reads index.md, picks up to four pages and answers from them. The wiki had 60 pages built from the same data; the vector store held 1,402 points. They graded 13 questions by hand, a single-tenant subset of 43. Their numbers:

Measure (their run) Hybrid RAG LLM wiki
Correct answers 13 of 13 12 of 13
Tokens per question, average 1,407 13,055
Latency, median 3.3 s 8.2 s
Latency, 95th percentile 24.3 s 27.3 s
LLM calls per question 2 4
Building the wiki None About 194,000 tokens, once

Adding one new source took 3,000 to 5,000 tokens in their setup. The two extra calls are by design: the wiki pipeline replaces search with two model steps, pick pages and then answer. On the one out-of-scope question the wiki was faster (1.6 s against 2.3 s), because the index step found nothing and stopped.

Two limits they state themselves: 13 questions is a small set, and they were mostly direct FAQ lookups, where RAG is strong. The case a wiki is built for, a question that needs several documents put together, was not in it.

Which is more accurate?

Neither by default. The wiki’s one miss in that evaluation is worth reading. Asked whether a policy could be transferred, it said yes and listed another tenant’s quotas, with a self-reported confidence of 0.9. The right answer for this tenant was no. The cause was at write time: pages were grouped by entity name rather than by tenant and entity, so two tenants’ policies were merged into one page. They found two more patterns:

  • A small model writing pages added noise (a stray line of Arabic and English in a Chinese answer). It was saved in the page, so every later answer repeated it.
  • Short paraphrased questions got one-line answers that dropped exceptions the RAG answer listed.

The wiki also reported higher average confidence (0.862 against 0.808) while scoring lower. Their conclusion: hybrid RAG stays the default, the wiki becomes an opt-in mode per request, and before switching, add a reranker, contextual retrieval, intent routing and tenant filters on the RAG side.

Two other results point the same way from different angles. Letta put LoCoMo conversation histories in files and gave a gpt-4o-mini agent file tools, grep and semantic search; it scored 74.0%, against the 68.5% Mem0 reported for its best graph variant. Letta’s reading is that how the agent searches (rewriting queries, searching again) matters more than the storage behind it. And a study of AGENTS.md files (arXiv 2602.11988) found that context files did not generally raise coding agents’ task success and raised inference cost by more than 20%. Repository overviews did not help; specific instructions were followed. Pre-written summaries are not free accuracy. What to load into context, and when, is the subject of our context engineering guide.

Which stays fresher?

RAG is current as soon as a document is indexed, and nobody rewrites anything. But it has no notion of “replaced”: the March decision and the June change sit in the index with equal standing, and the model sorts them out at answer time, every time.

A wiki is only as current as its last ingest. If sources arrive and nobody tells the agent to process them, the pages lag. In exchange, a contradiction is handled once, when the new source arrives, and the page says what changed and when. The gist’s lint step, which looks for contradictions, stale claims and orphan pages, only runs when someone asks for it.

Where does each stop scaling?

Karpathy says the index file “works surprisingly well at moderate scale (~100 sources, ~hundreds of pages)” without embedding-based retrieval. Past that he suggests proper search over the pages and names qmd, a local search engine for markdown with hybrid BM25 and vector search, LLM re-ranking, a CLI and an MCP server. In other words, a large wiki needs retrieval again; it retrieves pages instead of raw chunks.

RAG scales to collections nobody would want summarized: millions of chunks, documents that change hourly, archives searched once a year. Its limit runs the other way. Questions that need many documents at once get top-k fragments.

Boundaries are a wiki-specific risk. A model writing pages merges whatever looks alike unless the schema says otherwise. If your data has partitions (customers, tenants, teams that must not see each other’s information), put the partition in the page path or frontmatter and filter on it, or keep separate wikis.

When should you use RAG and when an LLM wiki?

Use RAG when:

  • the collection is large, or changes faster than anyone would re-ingest it;
  • questions are lookups with the answer in one place: FAQs, policies, docs;
  • answers must trace to the original wording, or be filtered per customer;
  • cost and latency per question matter (about a ninth of the wiki’s tokens in the evaluation above).

Use an LLM wiki when:

  • knowledge builds up over weeks: research, decisions, how systems fit together;
  • questions span several sources and similar ones come up again;
  • people need to read, and correct, what the agent believes;
  • several agents should start from the same conclusions instead of re-deriving them.

Our agent memory guide sets both next to memory files and memory APIs.

Can you combine RAG and an LLM wiki?

Yes, and most working setups do. The gist itself adds search once the wiki grows. Four patterns:

  • Search over the pages. Keep the wiki, and index its pages for keyword or vector search instead of relying on one index file.
  • Wiki first, raw sources second. Answer from the pages; when they do not cover the question, fall back to RAG over raw/ and file the answer as a new page, so the next person reads it.
  • Route by question. The evaluation above sends lookups to hybrid RAG and keeps the wiki as a per-request mode.
  • Lint on a schedule. Merged partitions and noise saved in pages are what a lint pass is for, if it runs.

Several tools in our LLM wiki GitHub roundup already pair a wiki with vector search. For memory that agents reach as an MCP server, see the best MCP memory servers.

Share the wiki with your other agents

A wiki pays off when the next session, and the next agent, starts from what was already worked out. That only happens if every agent reads the same pages.

Dexio is a hosted wiki that agents read and write over MCP. Claude Code, Codex, Cursor, Hermes Agent, OpenClaw, Claude and ChatGPT connect to one wiki of linked markdown pages, and you see what your agents know as a page graph. Raw files such as PDFs, decks and spreadsheets can sit beside the pages, and search_pages searches their text too. Every change records which agent made it. Dexio is open source (AGPL-3.0), free for one person, and $10 (Team) or $20 (Business) a member a month.

Dexio’s search matches every word of a question, best matches first. For semantic search across a large document collection, a RAG stack is the better pick; run it beside the wiki and have agents file what they conclude. To connect an agent that runs commands, tell it “Look at dexio.wiki and log me in.” Setup per agent: Claude Code, Hermes Agent and the rest in Guides.

FAQ

Does an LLM wiki replace RAG? Not on the evidence published so far. It replaces the synthesis RAG redoes on every question, not retrieval: once a wiki outgrows its index file, finding the right pages is a retrieval problem again.

Does an LLM wiki need embeddings? No. Karpathy’s version reads index.md and follows links. Larger wikis add search, sometimes with vectors.

Is NotebookLM RAG? The gist puts NotebookLM, ChatGPT file uploads and “most RAG systems” in the group that retrieves chunks at question time.

How many sources can an LLM wiki handle? About 100 sources and a few hundred pages with only an index file, per the gist. Beyond that, add search over the pages.

Sources

Ask a question

Ask anything about Dexio.

About
We reply by email.