cd /news/ai-agents/show-hn-zerikai-memory-local-code-me… · home › topics › ai-agents › article
[ARTICLE · art-142535] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Show HN: Zerikai Memory: local code-memory MCP server, now with Jev

Zerikai Memory, a local code-memory MCP server from developer KikeVen, added Jev, TypeSafe AI's System One judgment model, which scores per-passage relevance, evidence, contradiction and prompt-injection risk and returns calibrated probabilities instead of prose. Jev is off by default (ENABLE_JEV=false keeps behavior byte-identical) and is enabled by setting TYPESAFE_API_KEY and ENABLE_JEV=true in .env before restarting the server. The tool parses codebases with tree-sitter into a local ChromaDB vector store and serves context over a local STDIO MCP interface, and the maintainer develops on Windows while GitHub Actions verifies installation compiles on Windows, macOS and Linux.

read16 min views1 publishedSep 30, 2026
Show HN: Zerikai Memory: local code-memory MCP server, now with Jev
Image: Michielbdejong (auto-discovered)

⭐ Bookmark the project: If you use this tool, drop a star to save it to your GitHub profile and track new performance updates.

Never lose your AI context again.

zerikai_memory provides persistent, workspace-isolated memory for every IDE that is local-first, cost-aware, and instant. It uses deterministic Tree-Sitter code parsing indexing to capture entities and deep code descriptions like functions, classes, and docstrings into a local ChromaDB vector store. Accessed via a local MCP interface to slash token costs while maintaining high-resolution codebase mapping, it retrieves hyper-relevant context on query through L2 and Lexical re-indexing with strict source verification (Entity, File, Line Number, and L2). Designed to pair perfectly with low-cost DeepSeek APIs, it injects structured, highly precise local context instead of dumping raw, massive files, maximizing KV cache hits to radically reduce your active token costs.

💡Status: Active & Self-Contained.

This project is used daily and actively maintained by the author. Pull Requests and Issues are closed to keep maintenance overhead low. It is provided fully functional and ready for production use.

Platform Support Notice:

Note: This project is developed and actively maintained on Windows. GitHub Actions verifies that basic installation compiles across Windows, macOS, and Linux, but the maintainer cannot troubleshoot platform-specific runtime errors on Mac or Linux. Community pull requests fixing Mac/Linux bugs are highly welcome!

zerikai_memory now integrates Jev, TypeSafe AI's System One model — a fast, structured judgment engine that decides which retrieved passages actually answer your question and attaches a plain-text evidence report to every answer.

  • What it is: a second, narrow AI layer beside your LLM. Your LLM writes the answer; Jev judges the evidence it is built from. It returns calibrated probabilities — not prose.
  • What it does: per-passage relevance / evidence / contradiction / prompt-injection scoring, better passage ordering (replaces keyword rerank), and a plain-textAssessment / Evidence / Guidance report the agent can act on. Injection attempts are dropped.
  • Off by default: withENABLE_JEV=false behavior is byte-identical to before, and it is fully fail-open — if Jev is unavailable, the normal pipeline runs.
  • Activate it: setTYPESAFE_API_KEY andENABLE_JEV=true in.env , then restart the server.
  • Get the API key: TypeSafe early access →console.typesafe.ai .

📖 Full details — how the judgment layer works and every parameter: documentation/10-jev-judgment-layer.md

Every new chat session starts completely cold. When you switch contexts or open a new window:

  • Your AI Agent forgets every architectural decision, convention, and stack choice made over hours
  • You waste critical tokens and 10–15 minutes re-explaining the codebase setup in every single chat
  • Large raw file dumps inflate your token costs and shrink your available context window instantly
  • Switching IDEs (e.g., VS Code to Cursor) forces you to restart your conversation history from scratch

Zerikai Memory runs as a local STDIO MCP server between your IDE and your LLM. It parses your codebase using tree-sitter, indexes code entities into a local ChromaDB vector store, and injects highly relevant context snippets dynamically through natural language.

Your Codebase  →  tree-sitter (local parse)  →  ChromaDB (.brain/)
                                                      │
Your IDE       →  MCP Server (:stdio)        →        ▼
                                             Ollama / DeepSeek
                    │                      (auto-routed synthesis)
          ┌─────────┴──────────────┐
          │   4-Stage Pipeline     │
          │   L1  Vector Search    │  ChromaDB L2 distance matching
          │   L2  Lexical Re-rank  │  Keyword overlap boost on names
          │   L3  Auto-Routing     │  Ollama (free) vs. DeepSeek Cloud
          │   L4  LLM Synthesis    │  Answer + inline #file:line citations
          └────────────────────────┘
What gets taxed Without Zerikai With Zerikai
🔴 Monthly quota Re-explaining stack, decisions, and conventions every session Indexed once. Retrieved as compact snippets per query.
🟡 Context window Raw file dumps shrink the window available for code generation 1,000–1,200 token brief prefix*. Window stays wide open.
⚪ IDE switching Full re-explanation required in every new tool Shared zerikai_memory workspace .brain/ directory.

Tip: The project brief acts as a stable prefix. After your first query, DeepSeek caches it — making subsequent repeated queries 50–100× cheaper (rate depends on whether you are in peak or off-peak hours). See the DeepSeek Pricing section for current rates.

I have added a YouTube video walkthrough of the installation and setup process in the first step below. If you prefer text instructions, just follow along with the code snippets.

Watch the installation video #

Click the image below to watch a step-by-step walkthrough of the installation and setup process:

git clone https://github.com/your-username/zerikai_memory.git
cd zeriakai_memory

python -m venv venv
source venv/bin/activate  # Windows: .\venv\Scripts\activate

pip install -r requirements.txt

Remove the .example from the .env.example file in the root directory and rename it to .env:

Expand to view .env #

DEEPSEEK_API_KEY=your_deepseek_key_here

MEMORY_MODE=cloud

ENABLE_TOKEN_TRACKING=true

ENABLE_DEEPSEEK_PRO=false

DEEPSEEK_THINKING_BRIEF=disabled   # 9-section project brief generation
DEEPSEEK_THINKING_SCAN=disabled    # per-file indexing summaries
DEEPSEEK_THINKING_QUERY=disabled   # query_memory answer synthesis

DEEPSEEK_REASONING_EFFORT_BRIEF=low
DEEPSEEK_REASONING_EFFORT_SCAN=low
DEEPSEEK_REASONING_EFFORT_QUERY=low

QUERY_DISTANCE_THRESHOLD=1.0

SKIP_BARE_FILES=['.py', '.html', '.md', '.css']

ENABLE_LEXICAL_RERANK=true

LEXICAL_RERANK_WEIGHT=0.05

BRIEF_FETCH_CAP=20
python
python -c "from main import scan_workspace, query_memory; print('OK')"

You should see the startup banner followed by OK.

To stop your AI agent from ignoring the memory protocol, copy these directives into your IDE's agent rules profile (e.g., .cursorrules or system prompt guidelines):

IDE Rules in: agent_rules/ide_agent_rules.md

  • Universal-Brain First: The agentmust queryuniversal-brain before attempting raw file searches.
  • Source Discipline: Every answermust surface actualfile.py:line citations with zero fabrication.
  1. Press Ctrl+Shift+P →MCP: Add Local Server
  2. Choose STDIO
  3. Set command: C:\path\to\zerikai_memory\venv\Scripts\python.exe C:\path\to\zerikai_memory\main.py

Add to your claude_desktop_config.json profile:

{
  "mcpServers": {
    "universal-brain": {
      "command": "C:\\path\\to\\zerikai_memory\\venv\\Scripts\\python.exe",
      "args": ["C:\\path\\to\\zerikai_memory\\main.py"]
    }
  }
}

Works like .gitignore: one pattern per line. scan_workspace reads this file and skips matching paths.

Each project should have its own .memignore in its root directory. Forgetting to configure it before the first scan is the most common reason to use drop_memory.py and start fresh:

Examples of what to ignore: Expand to view

Sample .memignore #

.git/
node_modules/
venv/
__pycache__/
.brain/
dist/
build/

**/test/
**/tests/
.env
*.log
*.lock
*.pyc

Before running your first index scan, optimize your codebase's docstrings for vector search. Ask your AI Agent:

  • To install the embedding-docstring globally in your IDE and run it against your codebase to rewrite docstrings into a more embedding-friendly format.
  • You can find it in the

"Audit and optimize docstrings across this project using the embedding-docstring skill, respecting .memignore."

Requirement Why It Matters Target Impact
Explicit Tech Names Use "Uses Redis" instead of"key-value store" Embeddings match precise tokens, not abstract concepts.
Routing / Branches Document specific route paths and logical pivot options Ensures structural code matches are surfaceable.
Guarantees & Effects Explicitly state code idempotency, atomicity, or mutation side-effects Prevents agent generation from breaking runtime boundaries.

Simply instruct your IDE's active AI agent using natural language commands:

prefix queries with "universal-brain: <command>" to ensure they route through the MCP server and leverage your indexed memory:

  • Scan the workspace for the first time: "Set up memory for this project"
  • Ask a question: "What are the main architectural components of this project?"

Frequently used follow-ups:

  • After a code change: "Rescan the workspace and force a refresh of the project brief."
  • Save part of a chat: "Save the following context to memory: [your custom notes or constraints here]"
  • Ask how much have you used: "Get me a cost report for my memory usage so far."

See below for a full reference of available commands and their descriptions.

Upgrading zerikai_memory while keeping your already indexed workspaces is safe as long as the workspace path never changes.

Your workspace identity is a stable UUID derived from its absolute filesystem path (_derive_workspace_id in main.py). Because .brain/ is gitignored, a git pull never touches your indexed memory, project briefs, or workspace registry.

Upgrade in place — never copy, rename, clone, or move the project directory.

The whole point is to preserve your already-indexed .brain/ data — back it up, upgrade the code, restore it, and you get your old memory back exactly as it was. No re-indexing needed.



git pull


Do NOT rename the folder (e.g. zerikai_memory_old), clone into a differently-named folder, or move the project to a new parent directory. Any of these changes the path → generates a new UUID → orphans your old collection and brief. If you already did this, recover with merge_workspaces.

⚠️

Full step-by-step guide, backup/restore, and recovery instructions: documentation/09-upgrading.md

You never run these commands directly; your active AI agent executes them on your behalf.

Tool Description
init_workspace Registers a project folder, assigns a UUID, and creates a pending brief file. Idempotent; safe to run multiple times.
list_workspaces Lists all known workspaces that have a brief or stored memories.
resolve_workspace Resolves a workspace identifier (UUID, short-UUID, or display name) to its filesystem path.
merge_workspaces Consolidates duplicate workspace IDs into one. Irreversible.
debug_workspace_id Diagnostic tool; shows what workspace ID would be generated from a given path.
Tool Description
scan_workspace Starts a background scan. Returns immediately; use scan_status to track progress. Walks the directory, respects.memignore , saves all readable text files to persistent memory. Idempotent and self-cleaning. Concurrent (4 workers, batch writes).
scan_status Returns progress of a running or recently completed background scan: files scanned, entities indexed, errors, elapsed time, brief status.
save_to_memory Manually saves an architectural decision, fact, or technical note with an optional category tag.
list_memory Lists stored memories for a workspace, optionally filtered by category.
query_memory Retrieves relevant context via vector search and synthesises an answer via Ollama or DeepSeek (auto-routed). Returns the answer as plain text plus a trailing Sources: block offile:line citations with relevance scores (L2 distance or rerank).
get_brief Retrieves the current project brief from .brain/contexts/ .
update_brief Manually updates the markdown content of a project brief.
Tool Description
get_token_usage Returns DeepSeek API token usage and cost statistics.
get_cost_report Generates a cost breakdown by operation type. Prepends a live PEAK / OFF-PEAK banner showing currently active rates.
get_cache_stats Shows cache hit/miss rates by operation type.
purge_usage_data Deletes historical token tracking records.

When a workspace is scanned, Zerikai compiles a dense 1,000–1,200 token project brief across 9 locked components:

Section What It Captures
1. Overview Project domain, primary type, and functional scope.
2. Technical Stack Backend engines, databases, integrations, and core libraries.
3. Core Architecture Interactivity between frontend, backend, and processing layers.
4. Primary Conventions Local code styling, custom error handling, and validation schema rules.
5. Purpose Business logic problems solved and key underlying objectives.
6. Key Files Definitive app entry points, central routers, and specific domain tasks.
7. Dev & Testing Environment installation setups, execution triggers, and testing runs.
8. Data Flow Complete systemic request lifecycle tracing from gateway to database layer.
9. Future Roadmap Planned engineering steps and dangling TODO items parsed directly from code.

Adjust your operation profile via the MEMORY_MODE environment toggle to balance privacy, speed, and API costs.

💡 Deterministic First: All high-resolution code parsing (functions, classes, methods) is performed locally and deterministically using Tree-Sitter for $0 cost. The engines below are only used for text-file fallbacks and generating the architectural Project Brief.

Mode Analysis & Briefs Query Engine Total Cost Ideal Use Case
🟢 cloud DeepSeek DeepSeek Low Recommended. High-fidelity architectural briefs.
🟡 hybrid Ollama Ollama + DeepSeek Lowest Local privacy with cloud reasoning escalation.
🔴 local Ollama Ollama $0.00 100% air-gapped hardware-local tracking.
Key Default Description
DEEPSEEK_API_KEY Required Active API authorization key from platform.deepseek.com.
MEMORY_MODE cloud Sets target engines: choices include cloud ,hybrid , orlocal .
ENABLE_TOKEN_TRACKING true Calculates continuous usage and outputs summaries to SQLite.
QUERY_DISTANCE_THRESHOLD 1.5 Sets L2 vector distance cutoff limits. Lower inputs restrict matches.
ENABLE_LEXICAL_RERANK false Activates secondary hybrid reordering layer via keyword matching.
SKIP_BARE_FILES [] Extension list to bypass when tree-sitter finds zero valid code entities.

DeepSeek uses peak / off-peak pricing across all tiers. Off-peak rates are exactly half of peak rates.

⏰ Peak hours (UTC): 01:00–04:00 and 06:00–10:00, Monday–Friday only. Weekends are always off-peak.

Tier deepseek-flash input deepseek-flash output deepseek-flash cached v4-pro input v4-pro output v4-pro cached
Peak $0.30 $1.20 $0.006 $1.32 $3.96 $0.044
Off-peak $0.15 $0.60 $0.003 $0.66 $1.98 $0.022

The tool automatically resolves the correct tier at call time — no manual configuration needed.

Peak hours apply Monday–Friday only. Weekends are always off-peak regardless of time. Offsets shown for summer / daylight saving time (DST). In winter, US timezones shift 1 hour later; European zones shift 1 hour earlier — meaning off-peak windows shift accordingly.

⚠️

Region UTC offset (summer) Peak local time (Mon–Fri) ✅ Off-peak local time
EST (New York, Miami) UTC−5 8pm–11pm & 1am–5am 5am–8pm and 11pm–1am (+ all weekend)
CST (Chicago, Dallas) UTC−6 7pm–10pm & midnight–4am 4am–7pm and 10pm–midnight (+ all weekend)
PST (Los Angeles, Seattle) UTC−8 5pm–8pm & 10pm–2am 2am–5pm and 8pm–10pm (+ all weekend)
Ireland (Dublin) UTC+1 2am–5am & 7am–11am 11am–2am and 5am–7am (+ all weekend)
Spain (Madrid) UTC+2 3am–6am & 8am–noon noon–3am and 6am–8am (+ all weekend)
Germany (Berlin) UTC+2 3am–6am & 8am–noon noon–3am and 6am–8am (+ all weekend)
Norway (Oslo) UTC+2 3am–6am & 8am–noon noon–3am and 6am–8am (+ all weekend)

Key insight: For US users working standard hours on weekdays, most of the working day is already off-peak (cheaper). European users in GMT+2 zones benefit from off-peak pricing through most of the afternoon and evening — and all weekend queries cost even less. Saturday and Sunday queries are always billed at off-peak rates.

If you accidentally execute a workspace crawl before setting up your .memignore configurations, run the auxiliary wipe script to delete stale workspace data:

.\venv\Scripts\python.exe drop_memory.py "Workspace Name"

venv/bin/python drop_memory.py "Workspace Name"

Monitor server activity, runtime operations, and auto-routing logs inside .brain/server.log:

tail -f .brain/server.log

Get-Content .brain\server.log -Wait -Tail 30
  • All active vector spaces, tracking registries, and context details reside directly on your local machine.
  • Add .env and.brain/ explicitly to your global or project.gitignore patterns to prevent API keys and secure indexes from leaking to version control platforms.

To read more about the underlying design principles, architecture decisions, and future roadmap for Zerikai Memory, check out the insight article.

MIT License © Zerikai

🛠️ Support: This project is provided as-is for personal use.

── more in #ai-agents 4 stories · sorted by recency
── more on @zerikai memory 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-zerikai-memo…] indexed:0 read:16min 2026-09-30 · —