"The tedious part of maintaining a knowledge base is not the reading or the thinking — it's the bookkeeping." That is how Andrej Karpathy explains why people abandon their wikis: the maintenance burden grows faster than the value. The llm-wiki.md gist, published on 4 April 2026, is not a library, a benchmark or a model — it is a fifteen-hundred-word description of a pattern, meant to be pasted into your own agent and instantiated together with it. On 22 September we did exactly that in one of our production projects — AI Budget Assistant — and moved 24,616 words out of the project's main file in a single day. Below: what the pattern is, what it turned into on our side, and above all what still does not work in our implementation.
What Karpathy proposed
The usual experience of working with a language model and documents looks like RAG (Retrieval-Augmented Generation) lets an AI look up your own documents before answering, so its replies are grounded in your data instead of guesswork. Full definition →: you upload a collection of files, the model retrieves relevant A chunk is one of the fragments a document is split into, so an AI can retrieve and quote the relevant piece instead of reading the whole file. Full definition → for each question and generates an answer. It works, but the knowledge is rediscovered from scratch every time. Ask something that requires synthesizing five documents and the model has to find and piece the fragments together again. Nothing is built up. NotebookLM, chat file uploads and most RAG systems all work this way.
The gist proposes something different. Between you and the raw sources sits an An LLM wiki is a markdown knowledge base maintained by a language model: it reads new sources and weaves them into the pages that already exist, cross-references included, instead of searching everything again on every question. The pattern was described by Andrej Karpathy. Full definition → — a structured, interlinked collection of markdown files that the model builds and maintains incrementally. A new source is not merely indexed for later retrieval: the model reads it, extracts what matters and integrates it into the wiki that already exists — updating entity pages, revising topic summaries, noting where new data contradicts old claims. The knowledge is compiled once and then kept current, rather than re-derived on every query.
That is the whole difference: the wiki is a persistent, compounding artifact. The cross-references are already there. The contradictions have already been flagged. The synthesis already reflects everything that was read. You rarely write any of it yourself — the model does; your job is sourcing, direction and asking the right questions. Karpathy describes his own setup as the agent in one window and Obsidian in the other, with edits visible in real time. Obsidian is the IDE, the model is the programmer, the wiki is the codebase.
Three layers and three operations
| Layer | In the gist | On our side |
|---|---|---|
| Raw sources | An immutable collection of articles, PDFs and images | No such layer — the code itself is the source of truth and the wiki summarizes it |
| The wiki | A directory of markdown files the model owns | Domain hubs at the root, feature pages in a subfolder |
| The schema | A conventions-and-workflows file, e.g. CLAUDE.md or AGENTS.md | CLAUDE.md: repo-wide rules and the pointer to the index, no feature descriptions |
| The index | A catalog of every page with a one-line summary | A hub per domain, a page per feature, plus one rule on how to read a page |
| The log | A chronological record: ingests, queries, lint passes | The same three sections, one line per entry — the log is a search target, not a second wiki |
There are three operations as well. Ingest — a new source arrives, the model reads it, discusses the takeaways with you, writes a summary page, updates the index and every page it touches, and appends a line to the log; one source might touch ten to fifteen pages. Query — you ask a question, the model finds the relevant pages, reads them and answers with citations; and crucially, a good answer is filed back into the wiki as a new page, so your explorations compound the same way ingested sources do. Lint — a periodic health check: contradictions between pages, stale claims superseded by newer sources, orphan pages with no inbound links, important concepts lacking their own page, missing cross-references, data gaps.
Two special files sit apart from the rest. The index is content-oriented: a catalog of everything in the wiki, each page with a link and a one-line summary. The log is chronological: an append-only record of what happened and when. Karpathy notes that at around a hundred sources and a few hundred pages the index alone is enough, and An embedding turns the meaning of a text into numbers, which lets an AI find passages that mean something similar even when they use different words. Full definition → infrastructure is not needed at all.
How this differs from RAG and from ordinary documentation
From RAG — in where the synthesis happens. In RAG it happens at answer time and dies with the answer. In a wiki it happens at ingest time and stays on disk. The cost of a question drops not because retrieval got better, but because there is far less left to answer: half the work was done in advance.
From ordinary documentation — in who pays for maintenance. Docs rot not out of laziness but because the bookkeeping of cross-references is expensive while its value is spread thinly across the future. A model does not get bored, does not forget to update a reference, and can touch fifteen files in one pass. The cost of maintenance falls to near zero — and that is the only reason a wiki can survive at all.
There is one condition the gist states gently and we found to be hard: there has to be a loop, not a one-off act. We learned that the expensive way.
Why we took this on: 104k tokens before the first line of code
AI Budget Assistant is a monorepo: a NestJS An API is the interface two systems use to exchange data automatically, with no Excel exports and no copying by hand. Full definition →, an Expo mobile app, a Next.js admin dashboard, three chat bots and two shared packages. Its CLAUDE.md — the file an agent loads whole into every session — had grown to 77,100 words. That is roughly 104,000 A token is a fragment of text, usually a piece of a word, and it is the unit AI models measure length and charge by — bills are counted in tokens, not in questions. Full definition →, 187 top-level bullets, the largest of them 4,629 words.
The file was accurate. The end-of-task ritual appended to it everything learned along the way, and that worked honestly. What was wrong was the shape: an agent paid 104k tokens before it started working, and finding a single fact meant grepping a file with no navigation. A 4,629-word bullet is no longer documentation — it is four different subjects wearing one heading.
The wiki that lied for four months
Here is the interesting part: a docs/wiki directory already existed in that repository. Fifteen pages, written in May, in good shape. The one problem was that after May nothing was ever ingested into them again.
By September the wiki claimed 11 AI functions — there were 18; 8 languages — there are nine; 30 API modules — forty-eight. The health-trends file held exactly one datapoint, from 14 May.
This is not harmless. A stale wiki is worse than none: an agent reads it and reports the wrong number with confidence, and the person who received that number has no reason to double-check it. A missing page at least sends you to the code.
The very first reading pass after the relaunch found ten stale claims across two pages. The worst was the sync description: the page presented the generic queue as the mobile app's sync path, when two of that queue's key functions have zero call sites anywhere in the repository. Three of the same errors were live in CLAUDE.md and were fixed there too.
Our decision: CLAUDE.md becomes the schema
We wrote the design as its own document before doing the work — four decisions, each closing off a specific way to die.
One: CLAUDE.md stops being a content store and becomes the schema — the third layer from the gist. It keeps repo-wide rules, conventions, invariants belonging to no single feature, environment variables, the deploy procedure and the pointer to the index. Everything factual about a specific feature moves to a wiki page. One source of truth — because a wiki sitting next to a living, still-growing CLAUDE.md is exactly how the previous wiki died.
Two: migration is incremental, never a big-bang pass. A task that touches a feature moves that feature's description out of CLAUDE.md and onto its page as part of the same task. Re-homing 77,000 words in one sitting is mechanical work with a high chance of dropping nuance — and nuance is the whole value here.
Three: two levels — domain hubs and feature pages. The fifteen existing pages become hubs: what the domain is, where the entry points are, links to the feature pages under it. The granularity matches how work arrives: one task is usually one feature, so a task always has an obvious target page. A small thing, but it is what turns ingest from a decision into a reflex.
Four: hubs live at the root of the wiki directory (existing filenames kept — they are already linked from elsewhere), and feature pages in a subfolder.
The page template, and the section people open a page for
Six sections, one shape across the whole wiki. What this is — two or three sentences: what problem it solves, for whom. Entry points — files with paths, where to start reading. Key concepts — the mechanism, not a tutorial. Invariants — what must not break, stated as a rule, with the reason. Known gaps — what is deliberately not done and why, so nobody re-decides it. History — issue links: why it is this way.
Invariants is the section people open a page for, and it is precisely what the old wiki lacked. "Do not re-add a per-route sync-timestamp update", "resolve the server primary key before using an id as a foreign key" — those sentences existed in CLAUDE.md, but buried inside four-thousand-word bullets and reachable only by grep.
Two rules about being honest with the shape. A section that is genuinely empty is better omitted than filled with "none": an absent section honestly reads as "not yet examined", while the word "none" reads as a claim. And a page is not allowed to state a count: counts go stale silently, so when an audit finds a number on a page, the correct fix is to delete it, not to update it. This article is a dated artifact and does carry numbers; a wiki page has no such privilege.
Three rituals
| Operation | Our ritual | What it leaves behind |
|---|---|---|
| Ingest | Closing a task: the issue, the feature page, one line in the log | An updated page and one line in the ingests section |
| Query | Wiki before code; answer with the page path cited | A finding added to a page and a line in the queries section — even when no code changed |
| Machine lint | Two Python scripts weekly in CI, with no model call | A comment on one long-lived issue: dead links, orphans, pages the code has moved past |
| Reading lint | An in-session audit: two or three pages properly, not a skim of all | Corrected claims and a line in the lint-passes section |
Query deserves a separate note, because the first draft of our design did not have it — and without it the pattern is only half built. Ingest accumulates knowledge from changes. Without Query the wiki never starts accumulating knowledge from questions: a session that spends an hour proving why something behaves the way it does, and changes no code, leaves no trace at all. A diagnosis ending in "this is working correctly" is the worst case — a fix at least leaves a commit and an issue behind it, a clean bill of health leaves nothing, and it is exactly what the next session will re-investigate.
A second rule of the same ritual: answer with the page path cited. An answer with no citation is indistinguishable from one invented on the spot, and the person reading it can neither check the source nor improve it. And if the page disagreed with the code, that is a finding in itself: the page gets fixed in the same breath, rather than silently worked around.
Why our CI contains no model call at all
The lint is split in two by what can run where.
The machine-checkable half runs weekly in CI. The first script verifies that every link between pages resolves to a file that exists, that every cited path is on disk, and that no feature page is orphaned without a link from the index. The second compares a page's last commit date against the commit dates of the files it cites, and reports pages whose code has moved three or more commits ahead. The first exits non-zero; the second always exits zero — it is a report, not a gate. Nobody should be blocked from merging because a page is a week behind. Both land as a comment on one long-lived issue rather than a new issue each week, which would otherwise bury the repository in near-identical reports nobody reads.
No step of that workflow calls a model, and this is not asceticism. Claude Code on this project runs on a subscription, so there is no API key to put into repository secrets — and a job that needed one would simply never run. The reading half of the lint — contradictions between pages, missing cross-references, data gaps — is therefore a skill run inside a session, where the model is already paid for. The staleness report serves it as a priority order: a page with twenty commits behind it is where to look first; a page with three probably needs only a glance.
And one more rule derived from the earlier failure: read two or three pages properly rather than skimming all of them. A skim that concludes "looks fine" is the exact mechanism by which the previous wiki stayed wrong for four months.
Where we deliberately diverged from the gist
Three differences, all three downstream of our bottom layer not being a personal wiki's bottom layer.
We have no raw-sources layer. Karpathy's bottom layer is an immutable collection of articles, PDFs and images. Ours is the code itself, and the wiki summarizes it. That changes the nature of the error. In a personal wiki the source is static, and a page diverges from it only if it was written badly. In ours the source changes every day, so a page can become wrong without ever being edited. Hence the staleness report, which does not appear in the gist at all: we need a signal that the ground moved under a page.
We banned counts on pages. In a personal wiki a number is a fact from a source. In ours a number is almost always a count of something in the repository, and it is the first thing to go stale. This is a direct consequence of "11 AI functions" against 18: a page that never states a count cannot state one wrongly.
We are not building search. The gist suggests qmd, a local markdown search engine with hybrid BM25 and vector search. For now the index is enough: the wiki is read by path, the same way an agent reads code. No embeddings, no vector store — and we will revisit that when the index stops fitting in one glance, not before.
What the first day produced
One adoption task and two ordinary tasks with an ingest, and CLAUDE.md walked 77,100 → 63,916 → 59,394 → 52,484 words. The twelve heaviest bullets are drained. The largest of them, at 4,629 words, turned out to be four subjects under one heading and became four separate pages — which is itself a diagnosis of how the knowledge used to live.
Plus pages born not from migration but from ordinary work: a divergence of category ids between phone and server that made a filter return nothing, and a widget-refresh timer leaking across the mobile test suite's files and charging the failure to whichever suite happened to be running. Both pages exist only because the end-of-task ritual now leads into the wiki instead of into CLAUDE.md.
What still does not work in this scheme
The most honest section, and longer than we would like.
| Gap | Why it matters | What we will do |
|---|---|---|
| The log’s queries section is empty | The wiki accumulates only from changes, never from questions | Build filing an answer into a ritual instead of leaving it as a decision |
| 52,484 words still in CLAUDE.md | The rule "whoever touches it moves it" never moves a feature nobody touches | Either name the remainder what lives in the schema, or close the migration deliberately |
| The lint only checks repo-root paths | Paths cited relative to an app are not checked at all | Extend it only once it can be done without guessing a base directory |
| The staleness report counts commits, not substance | Three cosmetic commits look like three that broke the mechanism | Treat it as a priority order for a human, never as a verdict |
| The hubs have not been re-read since May | The top level of navigation is the artifact that already lied once | The next target is named in the log: the API page, 46 commits of schema behind |
| The health-trends file holds one datapoint from 14 May | Without a trend you cannot see whether the wiki is alive or dying | Append the weekly result to a file, not only to an issue comment |
| No measurement of the saving | Any savings figure quoted today would be invented | Measure the average volume read per session, before and after, on comparable tasks |
The log's Queries section is still empty. Ingest has fired six times, the reading audit once, Query zero times. As long as that holds, what we have is not a wiki in the gist's sense but well-organized project documentation that accumulates only from changes. The reason is plain: ingest is embedded in a mandatory end-of-task ritual, while filing an answer back is a separate decision that has to be made at the exact moment the question is already answered and you want to move on. Ritual beats intention — so what is needed is a ritual, not a reminder.
The migration will stop short of finishing, not because it finished. 52,484 words remain in CLAUDE.md, and the heaviest bullets are already out. From here each further move returns less for the same cost, and the rule "whoever touches it moves it" means features nobody touches will never move at all. That may well be correct — but then the remainder should honestly be called "what lives in the schema" rather than "the migration queue", or we will spend years counting as unfinished something we decided not to do.
Both scripts see less than they appear to. The path check only fires on paths anchored at the repository root, while pages legitimately cite paths relative to their own app — those citations are not checked at all. That is a deliberate compromise: a checker that guesses a base directory produces findings people stop trusting, which is worse than a narrow checker that is always right. But the price is that some references sit outside any control. The staleness report, meanwhile, remains a commit-count proxy: it does not know whether the substance changed, and three cosmetic commits look exactly like three that broke the mechanism a page describes.
The hubs are barely re-read. There are fifteen of them, written in May, and the reading audit has reached two. The next target is already named in the log — the API page, whose database schema has had 46 commits since the page was last touched. Until the hubs have been read against the code, the top level of navigation remains exactly the artifact that already lied once.
The health-trends file was designed as a time series and holds one datapoint from 14 May. While the weekly report goes out as an issue comment, there is physically nowhere for a trend to accumulate. And what shows whether a wiki is alive or dying is the trend, not a single snapshot: one week with five findings means nothing, five consecutive weeks with a rising count mean everything.
And the main one: we do not know what this saves. We have a separate model for agent cost, and it shows that the weight of context is one of the larger line items. But 104k tokens at the start of a session describes what was, not what is: how much the agent reads now, we have never measured. The honest answer to "what did this save in money" today is that we do not know, and that is a matter for measurement, not guesswork. The measurement is not hard: average volume read per session, before and after, on comparable tasks. We have not run it yet, and until we do, any savings figure in this article would be invented.
How to reproduce this in an evening
Paste the gist into your own agent and ask it to instantiate the pattern for your repository — that is literally what the document is for. It is deliberately abstract and ends by saying that its only job is to communicate the pattern and the agent can figure out the rest.
Then four things, each of which turned out to be mandatory for us.
Declare the schema in the file the agent always reads. Without a pointer to the index in the very first lines, a fresh session never learns the wiki exists and keeps appending to the old file. We wrote that into the log as its own line on launch day, because it is the one part of the scheme that cannot be deferred.
Fix one page template and put an invariants section in it. Everything else in the template can change later; invariants are the reason anyone opens a page before editing code.
Attach ingest to a ritual that is mandatory anyway. For us that is closing a task: issue, page, one line in the log. A ritual you have to remember does not get performed.
Put a cheap lint in place before the pages pile up. A couple of hundred lines of Python, no model, a weekly run. It will not find contradictions, but it will find dead links and orphans — exactly the defects nobody notices one at a time, until there are thirty of them.
One warning to close on. The moment a wiki starts lying looks precisely like the moment everything is fine: the pages are there, the links resolve, the agent answers confidently. The only difference is whether anyone has read those pages against the code in the last four months.
In summary
Karpathy's pattern does not solve retrieval; it solves accumulation. Knowledge is compiled once and kept current instead of being re-derived on every question. Inside a repository that becomes a simple rearrangement: the agent's main file becomes the schema, the knowledge moves onto pages, and the three operations become three rituals.
For us, the first day meant 24,616 fewer words in a file loaded into every session, and seventeen feature pages instead of list items. It has so far produced nothing from the second operation: not a single answered question filed back. And it has produced no measured savings figure — because we have not measured.
If there is one thing to take away, it is this: the artifact is almost never at fault. The previous wiki was well written and lied for four months because it had no loop. The pattern's value is not in the file layout but in making maintenance cheap enough that the loop can exist at all.
Frequently asked questions
- What is an LLM wiki in plain terms?
- It is a set of interlinked markdown files that the model builds and maintains for you. A new source is not merely indexed for later retrieval: the model reads it and integrates it into pages that already exist, updating summaries and cross-references. The knowledge is compiled once and kept current rather than re-derived on every question.
- How is an LLM wiki different from RAG?
- In where the synthesis happens. In RAG it happens at answer time and dies with the answer — the next question makes the model find and piece the fragments together again. In a wiki the synthesis happens at ingest time and stays on disk, along with the cross-references and the flagged contradictions. RAG optimizes retrieval; a wiki optimizes accumulation.
- Do you need a vector store or embeddings for this?
- Not at the start. Karpathy notes that at around a hundred sources and a few hundred pages an index file is enough: the model reads the index, picks the relevant pages and drills in. On our side the wiki is read by path, exactly the way an agent reads code — no embeddings, no vector store. The gist suggests adding a local markdown search engine only once the wiki grows past what the index can carry.
- Why is a stale wiki worse than no wiki at all?
- Because an agent reads it and reports the wrong number with confidence, and the person who received that number has no reason to double-check it. A missing page at least sends you to the code. Our own wiki, after four months without an update, claimed 11 AI functions against 18, 8 languages against nine, and 30 API modules against forty-eight — and all of those numbers were reported onwards as facts.
- Where do you start in your own repository?
- With four things that turned out to be mandatory for us. Declare the schema in the file the agent always reads — without a pointer to the index, a fresh session never learns the wiki exists. Fix one page template with an invariants section. Attach ingest to a ritual that is mandatory anyway, such as closing a task. And put a cheap, model-free lint in place before the pages pile up.