# How DeepWiki Works: Turning a Codebase into a Searchable Mental Model

> Source: <https://dev.to/shrsv/how-deepwiki-works-turning-a-codebase-into-a-searchable-mental-model-21jm>
> Published: 2026-10-06 19:47:14+00:00

*Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. [Star us](https://github.com/HexmosTech/LiveReview/) to help devs discover the project, give it a try, and share your feedback to help improve the product.*

An unfamiliar codebase does not feel difficult because it contains 100,000 lines of code.

It feels difficult because you do not know which 500 lines matter.

That is the problem DeepWiki is really solving.

When Cognition launched DeepWiki in May 2025, it described it as the public version of its internal wiki product. The basic interaction was  simple: take a GitHub URL, replace `github.com` with `deepwiki.com`, and get a generated wiki you can ask questions about. Cognition said more than 50,000 public repositories had already been indexed, including projects such as the Model Context Protocol and LangChain.

The interesting part is what has to happen between the URL and the answer.

DeepWiki cannot simply throw an entire repository into an LLM context window and say "explain this."

Instead, it builds several representations of the codebase:

```
                 Git repository
                       |
              clone / scan / filter
                       |
          +------------+------------+
          |                         |
    structural view            semantic index
    file tree + README          chunks + embeddings
          |                         |
          v                         v
   wiki structure                 FAISS
          |                         |
          +------------+------------+
                       |
                retrieved context
                       |
                       v
                     LLM
                       |
          +------------+------------+
          |                         |
      wiki pages                  answers
      + diagrams               + citations
```

The open DeepWiki-Open implementation gives us a useful way to inspect these mechanics. It should not be treated as Cognition's proprietary production source code; rather, it is an inspectable implementation of the same general product architecture.

There is a long history behind this idea.

In a 2012 observational study of 28 professional developers across seven companies, developers were found to use recurring comprehension strategies, often tried to avoid program comprehension as a task in itself, and frequently preferred face-to-face communication to documentation. The important point is that understanding software competes with the actual maintenance task the developer is trying to perform.

Suppose I ask:

How does authentication work in this repository?

There may be relevant code in:

```
src/auth/
api/middleware/
api/routes/
models/user.py
config/security.py
tests/auth/
```

The hard part is selecting the evidence.

A human engineer does this by building a mental model:

``` php
Request
  -> middleware
  -> token validation
  -> user lookup
  -> authorization
  -> handler
```

Only after this map exists does it become easy to dive into individual functions.

DeepWiki's architecture reflects the same decomposition.

There is a global representation of the repository that answers:

What exists, and how should I organize it?

And there is a retrieval representation that answers:

Which pieces are relevant to this particular question?

That distinction is the key to understanding the system.

Think of the first representation as a documentation outline and the second as a semantic search engine.

The documentation outline is hierarchical:

```
Repository
  |
  +-- Architecture
  |     +-- Request lifecycle
  |     +-- Data layer
  |     +-- Authentication
  |
  +-- Core modules
  |     +-- API
  |     +-- Services
  |     +-- Workers
  |
  +-- Deployment
        +-- Docker
        +-- CI/CD
```

The semantic index is much flatter:

``` php
chunk_001 -> src/api/router.py
chunk_002 -> src/auth/token.py
chunk_003 -> src/services/user.py
chunk_004 -> tests/auth/test_token.py
...
```

Each chunk gets converted into a vector.

This is important because the two structures optimize different tasks.

A hierarchy gives you coverage and navigation. A vector index gives you fast semantic lookup.

Cognition later pushed this idea another step by exposing DeepWiki through an MCP server. The official server provides programmatic operations such as `ask_question`, `read_wiki_contents`, and `read_wiki_structure`. In other words, the wiki is becoming an interface for software agents as well as a web page for humans.

This is a useful mental model:

DeepWiki is not primarily generating documentation. It is building an externalized representation of a codebase that humans and agents can query.

That is a more powerful way to think about it.

Now we can look at the first genuinely technical stage.

The open implementation starts by getting the repository locally. For remote repositories it can use Git and supports shallow cloning, which reduces unnecessary history and transfer cost.

It then walks the repository, filters files and directories, and turns source files into documents.

One practical constraint immediately appears: tokenization.

The implementation tracks token counts and keeps embedding chunks within a configured maximum, with `MAX_EMBEDDING_TOKENS` defaulting to 8192.

Why does this matter?

Consider a repository with 100,000 lines of code.

A rough estimate might be:

```
100,000 lines
x 5 to 10 tokens / line
-----------------------
500,000 to 1,000,000 tokens
```

Sending the whole repository to an LLM for every question would be wasteful even for very large-context models.

Instead, suppose the repository becomes 1,000 chunks.

Each chunk is embedded:

``` php
chunk -> embedding model -> vector
```

If the embedding dimension is 1,536, then storing the raw float32 vectors takes roughly:

```
1,000 vectors
x 1,536 dimensions
x 4 bytes
----------------
~6 MB
```

That is tiny compared with the original source tree.

The vectors are therefore an inexpensive reusable index.

The implementation also keeps metadata with each chunk, including things such as:

```
file_path
is_code
token_count
line information
```

The line tracking is important because retrieval is only useful if you can get back from:

```
"this chunk looks relevant"
```

to:

```
src/auth/token.py, lines 84-121
```

This is one of the places where the system stops looking like a simple "LLM wrapper" and starts looking like a conventional search system.

The historical connection here is semantic code search.

In 2019, a few academics published CodeSearchNet, a corpus containing about six million functions across six programming languages, together with expert relevance judgments for natural-language queries. They framed the central challenge exactly this way: natural language and programming-language expressions describe the same concepts using very different vocabulary.

DeepWiki sits on the same basic bridge:

```
"where are users authenticated?"
              |
              v
        query embedding
              |
              v
       nearest code chunks
              |
              v
        language model
```

This is perhaps the most important architectural decision.

A naive system might do:

``` php
repository -> LLM -> giant documentation
```

DeepWiki-Open instead performs an intermediate planning step.

It first reads the repository structure and README, then asks the LLM to produce a machine-readable wiki structure.

Conceptually:

```
<wiki_structure>
  <section title="Architecture">
    <page title="Request Lifecycle"
          importance="high"
          files="api/main.py,api/routers/..."/>
    <page title="Data Layer"
          importance="high"
          files="api/repository.py,..."/>
  </section>

  <section title="Deployment">
    ...
  </section>
</wiki_structure>
```

The actual implementation uses an XML structure containing sections and pages. A `WikiPage` carries information such as its title, importance, related pages, and the source file paths that should provide its context.

This is a form of coarse-to-fine generation:

``` php
repository
    |
    v
global structure
    |
    +---- page A -> relevant files -> generation
    |
    +---- page B -> relevant files -> generation
    |
    +---- page C -> relevant files -> generation
```

The advantage is that each individual generation problem becomes much smaller.

Suppose the model decides there should be 25 pages.

Rather than asking one model call to understand and document the entire repository, you can make roughly 25 focused generation problems.

The page prompt can then say, in effect:

```
You are writing the "Authentication" page.

Relevant files:
  api/auth.py
  middleware/security.py
  models/user.py

Explain:
  - how authentication starts
  - how credentials are validated
  - where identity is stored
  - important failure paths
```

That is a much more constrained task.

The open implementation goes further. Page prompts ask the model to identify the source files used for the explanation and impose formatting requirements for diagrams. The generated links are subsequently normalized into repository-specific URLs.

And there is a nice engineering detail hiding underneath all of this: the parser assumes the model can fail.

The structure parser strips formatting artifacts and has regex fallbacks when strict XML parsing fails, including cases where the model output gets truncated.

Once the repository has been indexed, answering a question becomes a classic RAG pipeline.

The open implementation's flow is roughly:

```
user query
    |
    v
query embedding
    |
    v
FAISS top-k retrieval
    |
    v
relevant code chunks
    |
    +---- conversation history
    |
    v
prompt builder
    |
    v
LLM
    |
    v
answer
```

The mathematics is straightforward.

Let the query embedding be:

```
q = [q1, q2, ..., qd]
```

and a code chunk embedding be:

```
x = [x1, x2, ..., xd]
```

A common similarity measure is cosine similarity:

```
cos(q, x) = (q dot x) / (||q|| ||x||)
```

The intuition is simply:

Do the query and this code chunk point in a similar direction in semantic space?

For example, a developer asks:

Where is the GitHub access token used?

A lexical search might miss a function called:

```
authenticate_repository(...)
```

because the words "access token" do not appear in the function name.

A semantic embedding can put concepts such as:

```
access token
OAuth credential
repository authentication
private clone
```

near one another.

Suppose there are 5,000 chunks and the embeddings have 1,536 dimensions.

A flat exact scan is roughly:

```
5,000 x 1,536
~= 7.7 million dimension comparisons
```

That is quite manageable in optimized native code. For much larger indexes, approximate nearest-neighbor structures can reduce the search work further.

The academic foundation for this style of architecture was laid out explicitly by Patrick Lewis and colleagues in the 2020 RAG paper. Their formulation combines a model's learned parametric memory with an external non-parametric memory represented by a dense vector index. The language model does the reasoning and generation; retrieval supplies the changing, inspectable evidence.

DeepWiki-Open follows this basic pattern, although its implementation is an application architecture rather than a direct reproduction of the RAG research model.

There is also a more expensive mode.

Its Deep Research path runs multiple iterations, with separate stages for planning, intermediate updates and final synthesis. In the inspected implementation, the loop can traverse the repository repeatedly, up to five iterations.

So the modes form something like:

``` php
Fast
  query -> retrieve -> answer

Deep Research
  query
    -> plan
    -> retrieve
    -> inspect
    -> refine
    -> retrieve again
    -> synthesize

Codemap
  query
    -> retrieve
    -> generate skeleton
    -> enrich sections
    -> ground citations
```

Codemap is particularly interesting.

The implementation first creates a skeleton, then enriches it with prose and diagrams. It does not simply trust the line numbers produced by the LLM. Instead, it can read the actual file, search for the cited snippet, and recover the real line range.

That creates a three-step chain:

```
LLM says:
  "this claim comes from foo.py"

        |
        v

system reads foo.py

        |
        v

system locates the snippet

        |
        v

citation points to actual lines
```

This is a small but important shift.

The model is allowed to propose evidence.

The deterministic program decides whether that evidence actually exists.

That principle generalizes far beyond documentation systems.

Once you view DeepWiki as an indexing and generation system, the operational design starts to make sense.

Wiki generation is too slow to treat as a normal HTTP request.

The open implementation therefore has a `TaskRegistry` and `WikiTask` state machine:

```
PENDING
   |
INDEXING
   |
DETERMINING_STRUCTURE
   |
GENERATING
   |
COMPLETED / FAILED
```

The frontend receives progress through Server-Sent Events.

This sounds like UI plumbing, but it is actually part of the product architecture.

A repository wiki might require:

```
clone
+ scan
+ embedding
+ structure generation
+ N page generations
+ post-processing
```

That can take long enough that a synchronous request would be the wrong abstraction.

The implementation also places explicit limits around concurrency.

The global wiki-task concurrency defaults to roughly half the available CPU cores, while the RAG layer separately uses a semaphore with a default concurrency of four.

That gives us a simple throughput model.

Suppose:

```
P = 24 pages
C = 8 concurrent page generations
t = 20 seconds average per page
```

Ignoring model-provider throttling and uneven page sizes:

```
T ~= ceil(P / C) x t
  ~= 3 x 20
  ~= 60 seconds
```

Serial execution would be:

```
24 x 20 = 480 seconds
```

So concurrency changes the economics of documentation generation dramatically.

But concurrency also creates failure modes.

The test suite in the open implementation contains two revealing cases:

```
test_submit_joins_active_task
test_page_failure_yields_placeholder_but_completes
```

The first prevents two requests for the same repository from kicking off duplicate work.

The second says that if one page fails, the entire wiki should still finish.

That is exactly the sort of detail you expect in a real asynchronous system. The expensive operation is the whole repository, so a single failed page should degrade one output rather than invalidate the entire job.

Caching is the other half of this economics.

The open implementation stores generated wikis in a local cache keyed by repository information such as host, owner, repository and language.

That creates a simple business equation:

```
cost per repository
  =
  clone cost
  + embedding cost
  + structure-generation cost
  + page-generation cost
```

The first two are largely preprocessing costs.

The expensive recurring component is usually generation:

```
generation cost
  ~ sum over pages of
      (retrieved input tokens + output tokens)
      x model price
```

Imagine:

```
25 pages
x 6 retrieved chunks/page
x 700 tokens/chunk
-------------------
105,000 retrieved input tokens
```

That is only an illustrative calculation. Actual context sizes, prompts, retries and outputs can move the number substantially.

The key architectural consequence is independent of the exact price:

You want to pay the expensive model cost once, then reuse the resulting representations many times.

That is why a wiki can make sense as a persistent artifact rather than regenerating an explanation on every question.

Cognition made this economics explicit in April 2026 when it said higher-quality DeepWiki generation modes would be usage-priced because these workflows consume meaningful compute, while the existing baseline generation experience would remain free.

The deeper idea is that DeepWiki is doing something closer to compilation than chat.

A compiler spends compute up front to create an intermediate representation so subsequent operations become cheaper.

DeepWiki spends compute up front to create:

``` php
repository
   -> structural representation
   -> semantic representation
   -> generated documentation
   -> cached knowledge
```

Then later questions can operate on those artifacts instead of repeatedly starting from raw source.

The easiest way to misunderstand DeepWiki is to think of it as:

"An LLM that writes documentation."

The more useful mental model is:

"A system that compiles a codebase into representations optimized for human and machine understanding."

The pieces line up:

``` php
Git repository
    |
    +--> file tree + README
    |        |
    |        +--> wiki structure
    |                 |
    |                 +--> generated pages
    |
    +--> chunks
             |
             +--> embeddings
                      |
                      +--> semantic retrieval
                               |
                               +--> RAG answers
                               +--> deep research
                               +--> codemaps
                               +--> citations
```

And this explains why the architecture has so many seemingly mundane components.

The embeddings matter because the repository is too large to read every time.

The wiki structure matters because search alone does not give you a coherent global model.

The task registry matters because generation takes time.

The cache matters because generation costs money.

The citation grounding matters because an LLM's claim that "this is line 84" is not evidence that line 84 actually contains the claim.

The parser fallbacks matter because models sometimes emit malformed structured output.

The whole system is therefore a combination of old ideas and new capabilities: information retrieval, software indexing, asynchronous job processing, caching, and program comprehension, with an LLM sitting in the middle as the component that turns selected code into explanations.

That may also be the larger pattern for AI software engineering.

The winning systems may be less about giving a model more context and more about building better representations of the context before the model ever sees it.

What would you add to DeepWiki's representation of a codebase before trusting it with a 500,000-line monorepo: dependency graphs, Git history, runtime traces, tests, production telemetry, or something else?

Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production reliable and secure without slowing you down.

I'm building **LiveReview**, a blast-radius aware AI code review built for your business-critical systems.

Instead of presenting every diff with equal emphasis, **LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.**

Spend code review effort where business risk is highest — not spread evenly across every diff.

⭐ Star it on GitHub: 

LiveReview is an AI code reviewer that scores every hunk of a diff by **blast radius**: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.

*LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.*

| The exact math, not a black box | Visualize blast radius at a glance | Every factor that feeds the score | 
|---|---|---|

**Here's the goal:**

**Click below to try LiveReview with your codebase:**
