I gave Claude Desktop a tax-free MCP memory layer A developer built zerikai_memory, an open-source MCP memory layer that runs locally behind Claude Desktop, using Tree-Sitter to parse source files and Markdown into a ChromaDB vector store. The tool aims to reduce token costs and context loss by providing cited, queryable project memory without LLM inference on the user's data. Most of us have felt it by now. The Context Tax. Slow token bleed just to re-establish what the AI already knew last session. A real dollar cost. But the problem started before AI tools existed. I once spent nearly a full week reverse-engineering a Django SaaS I was dropped into. Reading someone else's code. Mapping flows I had not written. Just to get to a place where I could build something new. Then came the AI boom. Same problem, different shape. I would return to a client project after a few months away, my own code, and stand there asking why I wrote a function that long. IDE agents helped. But every new session meant re-attaching files, re-explaining architecture, and burning 500 to 1,000 tokens before writing a single useful prompt. I tried managing it manually: NOTES.md file here.It was frustrating, but functional. Then the frontier model companies started repricing. I had an annual VS Code subscription active. I could see the math shifting under me. Before the price changes landed, I started researching memory layers. Not as a vague concept. As something I could build, own, and run locally. Something queryable. Something that traveled with the project. That is what zerikai memory https://github.com/KikeVen/zerikai memory became. It can run behind Claude Desktop via MCP. That is where it gets interesting. The more I learned about how LLMs hallucinate, the more I knew I had to stay as close to the source of truth as possible. The file scanner is entirely LLM-free. I reached for Tree-Sitter because it builds a Concrete Syntax Tree directly from source files. Full structural fidelity. Zero model inference touching your data. Right now it parses Python, JavaScript, TypeScript, HTML, and Markdown. More languages coming. What I did not expect was how well it handles Markdown. That turned out to be the real asset. When the scanner hits a .md or .mdx file, extract markdown in code indexer.py parses it into a CST and walks the tree recursively: Getting Started Installation Docker . ChromaDB vector store <- lives in .brain/, isolated per workspace | v Query triggers L1 - L2 - L3 - L4 L1: Vector search L2 distance, top-N results L2: Lexical re-rank boosts exact keyword + entity matches L3: Mode routing cloud / local / hybrid L4: Synthesis answer with inline file:line citations | v Your IDE gets clean, cited context. Not a file dump. The pipeline returns only the most relevant sections. The IDE or Claude Desktop session calls the memory via MCP. The context window stays lean. It does not dump entire files into the chat context. Precise source references back. Not raw context dumps. My API bill stays manageable. Once I realized I could drop research, reading logs, and saved notes into Markdown, scan them into memory, and query them through Claude Desktop, the workflow changed. Here is my deep research workflow: zerikai memory indexes it.The research workflow is one example. Any Claude Desktop workflow that benefits from persistent, queryable context can plug into the same pattern. Ten research results indexed. Sitting in memory. Accessible from any new chat session. I do not re-explain context. I do not re-paste 5,000-word documents. I just query. I can also ask the memory to save useful chunks from a live session under a custom title. A conversation that produced something worth keeping does not disappear when the terminal closes. A marketing brief built on indexed market data reads differently than one built from stale training data. A legal response built on precedent you actually looked up. A PRD that reflects real technical context. The memory layer is not just saving tokens. It changes what you can produce with them. For Markdown files, you still need an IDE to run the MCP server. Step 1: Clone and install git clone https://github.com/KikeVen/zerikai memory.git cd zerikai memory python -m venv .venv source .venv/bin/activate macOS / Linux .venv\Scripts\activate Windows pip install -r requirements.txt Verify python -c "from main import scan workspace, query memory; print 'OK' " Edit the .env file in the clone directory. Point it at your preferred LLM in Ollama or DeepSeek requires an API key . If there are files or directories you do not want scanned in the workspace you want indexed, add a .memignore file to the project root and list them there. The scanner skips them. Step 2: Add it to your IDE/coding CLI MCP config { "mcpServers": { "universal-brain": { "command": "/absolute/path/to/zerikai memory/.venv/bin/python", "args": "/absolute/path/to/zerikai memory/main.py" } } } Every path must be absolute. Relative paths cause silent startup failures. No error message. You will spend an hour debugging nothing. Step 3: Scan your workspace Open the VS Code command palette or IDE/coding CLI . Call universal-brain . Scan the workspace for the first time: "Set up memory for this project" It runs in the background. Poll with scan status to track progress. Once complete, the codebase is indexed. The project brief lives in .brain/contexts/ . For a full list of tool commands, visit docs https://github.com/KikeVen/zerikai memory mcp-tools-reference Step 4: Query it query memory: where does data flow after the payments endpoint? That is it. Visit zerikai memory https://github.com/KikeVen/zerikai memory on GitHub to install it.