Blog
Blog

Why Karpathy’s LLM Wiki makes sense for AI agents

A plain explanation of Karpathy’s LLM Wiki, how agents can navigate its knowledge, and why sources, updates, and human review still matter.

By Mrigesh Parashar
An index entry leads through a project decision to its supporting source while unrelated pages remain outside the selected route.

An agent can read every note from a project and still miss what matters. The decision might be in a meeting transcript, the exception in a later message, and the reason in a document nobody thought to attach.

Andrej Karpathy’s LLM Wiki is a compelling way to approach that problem. It gives useful understanding a place to survive between questions. The interesting part is what this changes about the next task.

The short answer

A maintained knowledge base can give an agent a useful starting point: what we currently believe, how the pieces relate, and where to check the evidence.

My reading of the design is simple: some of the work of answering tomorrow’s question can happen when we understand today’s source. Whether that pays off depends on how often the knowledge is reused and how well it is maintained.

What Karpathy proposes

Karpathy’s LLM Wiki idea file describes three layers: preserved source material, a linked Markdown wiki maintained by an LLM, and instructions for how to maintain it.

The workflow includes adding sources, querying the wiki, and checking its health. An index helps locate pages; a log records changes. Useful answers can become part of the wiki.

It is a proposal for a workflow, with implementation choices left to the reader. It is not a benchmark demonstrating that every agent will become more accurate.

Why an agent can find a useful path

Consider this fictional question:

Can we open the Atlas beta on Friday?

The project has these records:

RecordWhat it says
Monday planning meetingFriday is the target
Wednesday customer callExport must work before the customer can participate
Thursday test resultExport loses one attachment type
Thursday launch decisionKeep the beta closed until the export issue is resolved

A search for “Friday beta” could find the target date without surfacing the export dependency. A useful answer needs to connect the launch decision to the test result.

Here is how I would structure a small wiki for this example.

The project entry says: “Atlas beta — launch blocked on export; see launch decision.” The decision page names the condition and links to the export issue. The issue page links to the failing test and its date.

An agent has a route to investigate:

  1. Find the Atlas entry.
  2. Read the current launch decision.
  3. Follow the export dependency.
  4. Check the test evidence.
  5. Answer with the condition still attached.

That route is helpful because each step says why the next record matters. A filename such as “Thursday notes” leaves that relationship for the reader to reconstruct.

Anthropic’s context engineering guidance describes agents discovering context progressively through file names, metadata, search, and selective reads. It also describes retaining notes outside the active conversation. Those mechanisms help explain how an agent could navigate this example; they do not prove that our hypothetical wiki would retrieve the right answer every time.

The agent still needs tools that can read the files, permission to use them, and instructions to follow the evidence. A folder full of Markdown does not load itself into a model.

Why the saved explanation matters

Suppose someone asks a different question the next morning:

What should we tell the customer?

The same export dependency matters, even though the question contains neither “Friday” nor “beta.”

In our example, the project page already identifies the unresolved condition. The agent can use it to draft a careful update: the opening date remains unconfirmed because the export check has not passed.

This is what makes the design appealing to me. People ask about the same project from different angles. A useful explanation of its current state can serve more than one question.

It can also make review more precise. If the project page says the customer needs every attachment format, a reviewer can ask which source supports “every.” Correcting that sentence has a clear location and scope.

There is a tradeoff. Preparing and checking that page costs time and model usage. For a document you will read once, answering directly may be enough. For a project revisited throughout a month, maintained context is more plausible as an investment. That is a design judgment, not a measured cost saving.

Longer context windows do not settle the question. The Lost in the Middle study found that the position of relevant information affected performance on the models and tasks tested. It is evidence that fitting text into context and using it successfully are different problems. It is not a scorecard for every current model.

The maintenance rule I would insist on

Return to Atlas. A new test now shows that export works.

What should change?

The issue page can record the passing test. The launch page still needs an explicit decision about opening. A technical blocker disappearing does not, by itself, authorize a customer announcement.

I would keep those as separate statements:

  • Export passed the latest test.
  • The launch owner has not yet confirmed the opening date.

That distinction is easy to lose during summarization. It is also exactly the distinction an agent needs before drafting or taking an action.

A useful update process should preserve the old evidence, identify the changed claim, and show what remains unresolved. For consequential decisions, a person should review the proposed update before it becomes the record others rely on.

The failure mode deserves attention: one mistaken summary can be read by several later tasks. Reuse can spread an error as readily as a correction.

I would evaluate a wiki with questions whose answers are known. Include an outdated decision, a missing owner, and two sources that disagree. Check whether the agent finds the right evidence and admits the gaps. A tidy graph is not that test.

Where the idea has limits

An LLM Wiki still needs retrieval. An index, links, keyword search, and vector search can all help find material at different scales.

Nor is advance synthesis unique to this pattern. The GraphRAG paper describes building a graph and preparing community summaries before answering broad questions. That is a different architecture, but it is a useful reason to avoid saying that all retrieval systems start from nothing.

I would begin with this approach for a bounded project whose questions recur. I would be more cautious when information changes constantly, access differs by person, or an incorrect summary could carry a serious consequence. Those cases need stronger update, access, and review controls.

The practical starting point can be small. Our guide to organizing notes for later AI use shows how to create one current project note. Our working-memory guide covers bringing the relevant part into an agent’s task.

Start with a question you ask repeatedly. Build a page that helps answer it, preserve the evidence, and check what happens when the answer changes.

That is where a knowledge base earns its place in the work.

Frequently asked questions

Does the wiki become part of the model?

The files remain outside the model. A tool must retrieve or read the relevant content into a request. Writing a page does not train the model.

Does this replace RAG?

It can be combined with retrieval-augmented generation. The choices of what to store, what to summarize, and how to retrieve it are related but separate.

Is Markdown the reason it works?

Readable files make the material convenient to inspect and edit. In the example above, the more important design choices are the explicit dependency, current decision, and link to evidence.

Sources

Accessed October 3, 2026.