How DeepWiki Works: Turning a Codebase into a Searchable Mental Model Developer Shrijith Venkatramana explains the architecture behind DeepWiki, Cognition's tool that generates searchable wikis from GitHub repositories by replacing github.com with deepwiki.com. The system builds two separate representations of a codebase — a hierarchical structural view for wiki organization and a flat semantic embedding index in FAISS for question answering — rather than feeding an entire repository into an LLM context window. Cognition said more than 50,000 public repositories had been indexed at launch, including the Model Context Protocol and LangChain. Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us https://github.com/HexmosTech/LiveReview/ to help devs discover the project, give it a try, and share your feedback to help improve the product. An unfamiliar codebase does not feel difficult because it contains 100,000 lines of code. It feels difficult because you do not know which 500 lines matter. That is the problem DeepWiki is really solving. When Cognition launched DeepWiki in May 2025, it described it as the public version of its internal wiki product. The basic interaction was simple: take a GitHub URL, replace github.com with deepwiki.com , and get a generated wiki you can ask questions about. Cognition said more than 50,000 public repositories had already been indexed, including projects such as the Model Context Protocol and LangChain. The interesting part is what has to happen between the URL and the answer. DeepWiki cannot simply throw an entire repository into an LLM context window and say "explain this." Instead, it builds several representations of the codebase: Git repository | clone / scan / filter | +------------+------------+ | | structural view semantic index file tree + README chunks + embeddings | | v v wiki structure FAISS | | +------------+------------+ | retrieved context | v LLM | +------------+------------+ | | wiki pages answers + diagrams + citations The open DeepWiki-Open implementation gives us a useful way to inspect these mechanics. It should not be treated as Cognition's proprietary production source code; rather, it is an inspectable implementation of the same general product architecture. There is a long history behind this idea. In a 2012 observational study of 28 professional developers across seven companies, developers were found to use recurring comprehension strategies, often tried to avoid program comprehension as a task in itself, and frequently preferred face-to-face communication to documentation. The important point is that understanding software competes with the actual maintenance task the developer is trying to perform. Suppose I ask: How does authentication work in this repository? There may be relevant code in: src/auth/ api/middleware/ api/routes/ models/user.py config/security.py tests/auth/ The hard part is selecting the evidence. A human engineer does this by building a mental model: php Request - middleware - token validation - user lookup - authorization - handler Only after this map exists does it become easy to dive into individual functions. DeepWiki's architecture reflects the same decomposition. There is a global representation of the repository that answers: What exists, and how should I organize it? And there is a retrieval representation that answers: Which pieces are relevant to this particular question? That distinction is the key to understanding the system. Think of the first representation as a documentation outline and the second as a semantic search engine. The documentation outline is hierarchical: Repository | +-- Architecture | +-- Request lifecycle | +-- Data layer | +-- Authentication | +-- Core modules | +-- API | +-- Services | +-- Workers | +-- Deployment +-- Docker +-- CI/CD The semantic index is much flatter: php chunk 001 - src/api/router.py chunk 002 - src/auth/token.py chunk 003 - src/services/user.py chunk 004 - tests/auth/test token.py ... Each chunk gets converted into a vector. This is important because the two structures optimize different tasks. A hierarchy gives you coverage and navigation. A vector index gives you fast semantic lookup. Cognition later pushed this idea another step by exposing DeepWiki through an MCP server. The official server provides programmatic operations such as ask question , read wiki contents , and read wiki structure . In other words, the wiki is becoming an interface for software agents as well as a web page for humans. This is a useful mental model: DeepWiki is not primarily generating documentation. It is building an externalized representation of a codebase that humans and agents can query. That is a more powerful way to think about it. Now we can look at the first genuinely technical stage. The open implementation starts by getting the repository locally. For remote repositories it can use Git and supports shallow cloning, which reduces unnecessary history and transfer cost. It then walks the repository, filters files and directories, and turns source files into documents. One practical constraint immediately appears: tokenization. The implementation tracks token counts and keeps embedding chunks within a configured maximum, with MAX EMBEDDING TOKENS defaulting to 8192. Why does this matter? Consider a repository with 100,000 lines of code. A rough estimate might be: 100,000 lines x 5 to 10 tokens / line ----------------------- 500,000 to 1,000,000 tokens Sending the whole repository to an LLM for every question would be wasteful even for very large-context models. Instead, suppose the repository becomes 1,000 chunks. Each chunk is embedded: php chunk - embedding model - vector If the embedding dimension is 1,536, then storing the raw float32 vectors takes roughly: 1,000 vectors x 1,536 dimensions x 4 bytes ---------------- ~6 MB That is tiny compared with the original source tree. The vectors are therefore an inexpensive reusable index. The implementation also keeps metadata with each chunk, including things such as: file path is code token count line information The line tracking is important because retrieval is only useful if you can get back from: "this chunk looks relevant" to: src/auth/token.py, lines 84-121 This is one of the places where the system stops looking like a simple "LLM wrapper" and starts looking like a conventional search system. The historical connection here is semantic code search. In 2019, a few academics published CodeSearchNet, a corpus containing about six million functions across six programming languages, together with expert relevance judgments for natural-language queries. They framed the central challenge exactly this way: natural language and programming-language expressions describe the same concepts using very different vocabulary. DeepWiki sits on the same basic bridge: "where are users authenticated?" | v query embedding | v nearest code chunks | v language model This is perhaps the most important architectural decision. A naive system might do: php repository - LLM - giant documentation DeepWiki-Open instead performs an intermediate planning step. It first reads the repository structure and README, then asks the LLM to produce a machine-readable wiki structure. Conceptually: