Show HN: Zerikai Memory: local code-memory MCP server, now with Jev Zerikai Memory, a local code-memory MCP server from developer KikeVen, added Jev, TypeSafe AI's System One judgment model, which scores per-passage relevance, evidence, contradiction and prompt-injection risk and returns calibrated probabilities instead of prose. Jev is off by default (ENABLE_JEV=false keeps behavior byte-identical) and is enabled by setting TYPESAFE_API_KEY and ENABLE_JEV=true in .env before restarting the server. The tool parses codebases with tree-sitter into a local ChromaDB vector store and serves context over a local STDIO MCP interface, and the maintainer develops on Windows while GitHub Actions verifies installation compiles on Windows, macOS and Linux. ⭐ Bookmark the project: If you use this tool, drop a star to save it to your GitHub profile and track new performance updates. Never lose your AI context again. zerikai memory provides persistent, workspace-isolated memory for every IDE that is local-first, cost-aware, and instant. It uses deterministic Tree-Sitter code parsing indexing to capture entities and deep code descriptions like functions, classes, and docstrings into a local ChromaDB vector store. Accessed via a local MCP interface to slash token costs while maintaining high-resolution codebase mapping, it retrieves hyper-relevant context on query through L2 and Lexical re-indexing with strict source verification Entity, File, Line Number, and L2 . Designed to pair perfectly with low-cost DeepSeek APIs, it injects structured, highly precise local context instead of dumping raw, massive files, maximizing KV cache hits to radically reduce your active token costs. πŸ’‘ Status: Active & Self-Contained. This project is used daily and actively maintained by the author. Pull Requests and Issues are closed to keep maintenance overhead low. It is provided fully functional and ready for production use. Platform Support Notice: Note : This project is developed and actively maintained on Windows . GitHub Actions verifies that basic installation compiles across Windows, macOS, and Linux, but the maintainer cannot troubleshoot platform-specific runtime errors on Mac or Linux. Community pull requests fixing Mac/Linux bugs are highly welcome zerikai memory now integrates Jev , TypeSafe AI's System One model β€” a fast, structured judgment engine that decides which retrieved passages actually answer your question and attaches a plain-text evidence report to every answer. - What it is: a second, narrow AI layer beside your LLM. Your LLM writes the answer; Jev judges the evidence it is built from. It returns calibrated probabilities β€” not prose. - What it does: per-passage relevance / evidence / contradiction / prompt-injection scoring, better passage ordering replaces keyword rerank , and a plain-text Assessment / Evidence / Guidance report the agent can act on. Injection attempts are dropped. - Off by default: with ENABLE JEV=false behavior is byte-identical to before, and it is fully fail-open β€” if Jev is unavailable, the normal pipeline runs. - Activate it: set TYPESAFE API KEY and ENABLE JEV=true in .env , then restart the server. - Get the API key: TypeSafe early access β†’ console.typesafe.ai https://console.typesafe.ai . πŸ“– Full details β€” how the judgment layer works and every parameter: documentation/10-jev-judgment-layer.md https://github.com/KikeVen/zerikai memory/blob/main/documentation/10-jev-judgment-layer.md Every new chat session starts completely cold. When you switch contexts or open a new window: - Your AI Agent forgets every architectural decision, convention, and stack choice made over hours - You waste critical tokens and 10–15 minutes re-explaining the codebase setup in every single chat - Large raw file dumps inflate your token costs and shrink your available context window instantly - Switching IDEs e.g., VS Code to Cursor forces you to restart your conversation history from scratch Zerikai Memory runs as a local STDIO MCP server between your IDE and your LLM. It parses your codebase using tree-sitter, indexes code entities into a local ChromaDB vector store, and injects highly relevant context snippets dynamically through natural language. Your Codebase β†’ tree-sitter local parse β†’ ChromaDB .brain/ β”‚ Your IDE β†’ MCP Server :stdio β†’ β–Ό Ollama / DeepSeek β”‚ auto-routed synthesis β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ 4-Stage Pipeline β”‚ β”‚ L1 Vector Search β”‚ ChromaDB L2 distance matching β”‚ L2 Lexical Re-rank β”‚ Keyword overlap boost on names β”‚ L3 Auto-Routing β”‚ Ollama free vs. DeepSeek Cloud β”‚ L4 LLM Synthesis β”‚ Answer + inline file:line citations β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ | What gets taxed | Without Zerikai | With Zerikai | |---|---|---| | πŸ”΄ Monthly quota | Re-explaining stack, decisions, and conventions every session | Indexed once. Retrieved as compact snippets per query. | | 🟑 Context window | Raw file dumps shrink the window available for code generation | 1,000–1,200 token brief prefix . Window stays wide open. | | βšͺ IDE switching | Full re-explanation required in every new tool | Shared zerikai memory workspace .brain/ directory. | Tip: The project brief acts as a stable prefix. After your first query, DeepSeek caches it β€” making subsequent repeated queries 50–100Γ— cheaper rate depends on whether you are in peak or off-peak hours . See the DeepSeek Pricing deepseek-pricing section for current rates. I have added a YouTube video walkthrough of the installation and setup process in the first step below. If you prefer text instructions, just follow along with the code snippets. Watch the installation video Click the image below to watch a step-by-step walkthrough of the installation and setup process: git clone https://github.com/your-username/zerikai memory.git cd zeriakai memory Create and activate a virtual environment Python 3.11+ python -m venv venv source venv/bin/activate Windows: .\venv\Scripts\activate pip install -r requirements.txt Remove the .example from the .env.example file in the root directory and rename it to .env : Expand to view .env DEEPSEEK API KEY=your deepseek key here Memory Mode controls which LLM is used for operations: - "cloud": Use DeepSeek for all operations scan, brief, queries - highest quality, tracked usage - "hybrid": Use Ollama for file scanning, DeepSeek for briefs and escalated queries - "local": Use Ollama for everything free, but lower quality briefs MEMORY MODE=cloud Enable token tracking and cost reporting SQLite database at .brain/token usage.db Set to "false" to disable tracking ENABLE TOKEN TRACKING=true Enable deepseek-v4-pro for complex architectural queries design, architecture, tradeoffs v4-pro is ~4Γ— more expensive than deepseek-flash. See DeepSeek Pricing section below for current peak/off-peak rates. Recommended: keep this "false" unless you need maximum reasoning capability ENABLE DEEPSEEK PRO=false DeepSeek thinking mode chain-of-thought . Docs: https://api-docs.deepseek.com/guides/thinking mode/ Thinking is ON by default at effort "high" when no parameter is sent. Options per path: enabled | disabled. "disabled" is right for extractive work. NOTE: briefs and query synthesis are separate pipelines with separate toggles. DEEPSEEK THINKING BRIEF=disabled 9-section project brief generation DEEPSEEK THINKING SCAN=disabled per-file indexing summaries DEEPSEEK THINKING QUERY=disabled query memory answer synthesis Reasoning effort per path when that path is "enabled": low | high | max. Ignored when the matching path is "disabled". "low" is the cheapest enabled tier. DEEPSEEK REASONING EFFORT BRIEF=low DEEPSEEK REASONING EFFORT SCAN=low DEEPSEEK REASONING EFFORT QUERY=low Semantic search relevance cutoff for query memory L2 distance . Lower = stricter. Watch "best dist=X.XX" in server.log to calibrate. Typical: <0.8 strong match, 0.8-1.5 related, 1.5 noise. QUERY DISTANCE THRESHOLD=1.0 File extensions to skip during scanning when tree-sitter produces zero entities no functions, classes, headings, semantic HTML elements, etc. . Saves API calls on bare config files, trivial templates, empty CSS, etc. Format: '.py', '.html', '.md', '.css' Default: empty β€” no extensions skipped, all fall through to LLM . SKIP BARE FILES= '.py', '.html', '.md', '.css' Enable lexical re-ranking in query memory. When true, results passing the distance threshold are reordered by a weighted combination of semantic distance and keyword overlap in entity name and docstring text. Nothing is dropped β€” pure reorder. Default: false existing pure-semantic behaviour preserved . ENABLE LEXICAL RERANK=true Weight applied per keyword hit during lexical re-ranking. The 1/dist spread across the valid-hit band 0.85–0.98 is ~0.156. Keep this value below that spread to avoid keyword hits overriding a genuinely closer semantic result. Recommended starting point: 0.05 one hit = +0.05, two hits = +0.10 . LEXICAL RERANK WEIGHT=0.05 Candidate pool per section for project-brief synthesis. Each section queries ChromaDB, re-ranks locally, then trims to a per-section cap 20/25/30 . Decoupled from FETCH CAP query-only so a tight query pool doesn't starve the brief. Default: 20. BRIEF FETCH CAP=20 python python -c "from main import scan workspace, query memory; print 'OK' " You should see the startup banner followed by OK . To stop your AI agent from ignoring the memory protocol, copy these directives into your IDE's agent rules profile e.g., .cursorrules or system prompt guidelines : IDE Rules in: agent rules/ide agent rules.md https://github.com/KikeVen/zerikai memory/blob/main/agent rules/ide agent rules.md - Universal-Brain First: The agent must query universal-brain before attempting raw file searches. - Source Discipline: Every answer must surface actual file.py:line citations with zero fabrication. 1. Press Ctrl+Shift+P β†’ MCP: Add Local Server 2. Choose STDIO 3. Set command: C:\path\to\zerikai memory\venv\Scripts\python.exe C:\path\to\zerikai memory\main.py Add to your claude desktop config.json profile: { "mcpServers": { "universal-brain": { "command": "C:\\path\\to\\zerikai memory\\venv\\Scripts\\python.exe", "args": "C:\\path\\to\\zerikai memory\\main.py" } } } Works like .gitignore : one pattern per line. scan workspace reads this file and skips matching paths. Each project should have its own .memignore in its root directory. Forgetting to configure it before the first scan is the most common reason to use drop memory.py and start fresh: Examples of what to ignore: Expand to view Sample .memignore Directories trailing slash required .git/ node modules/ venv/ pycache / .brain/ dist/ build/ File/Folder patterns /test/ /tests/ .env .log .lock .pyc Before running your first index scan, optimize your codebase's docstrings for vector search. Ask your AI Agent: - To install the embedding-docstring globally in your IDE and run it against your codebase to rewrite docstrings into a more embedding-friendly format. - You can find it in the embedding-docstring skill guide https://github.com/KikeVen/zerikai memory/blob/main/embedding-docstring/SKILL.md . - You can find it in the "Audit and optimize docstrings across this project using the embedding-docstring skill, respecting .memignore." | Requirement | Why It Matters | Target Impact | |---|---|---| | Explicit Tech Names | Use "Uses Redis" instead of "key-value store" | Embeddings match precise tokens, not abstract concepts. | | Routing / Branches | Document specific route paths and logical pivot options | Ensures structural code matches are surfaceable. | | Guarantees & Effects | Explicitly state code idempotency, atomicity, or mutation side-effects | Prevents agent generation from breaking runtime boundaries. | Simply instruct your IDE's active AI agent using natural language commands: prefix queries with "universal-brain: