ContextMemory – Markdown memory for your llama.cpp/vLLM server Kortexio released ContextMemory, an open-source self-hosted memory gateway that sits in front of any OpenAI-compatible /v1 engine such as llama.cpp, vLLM, Ollama, LM Studio or OpenAI and stores session memory as editable markdown files rather than a vector database. The gateway authenticates tenants, injects a session wiki plus history, runs an agentic tool loop with sandbox and MCP tools, applies skills, guardrails, validators and optional human-in-the-loop confirmation, and returns a standard OpenAI-shaped chat.completions response with streaming support. It is installed via git clone and docker compose with llama.cpp or vLLM overrides, and its CI runs an end-to-end check against a real llama-server on every push. Try it 3 commands try-it-in-three-commands · Engines https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md engines · Docs https://github.com/Kortexio/ContextMemory/blob/main/docs/README.md · vs Mem0 / Zep / Letta https://github.com/Kortexio/ContextMemory/blob/main/docs/compare.md Self-hosted memory gateway for your llama.cpp / vLLM server. Put one OpenAI-compatible /v1 URL in front of your engine. Your client sends only the new message; the gateway keeps session memory as markdown you can open, edit, and diff . No vector DB, no client rewrite. git clone https://github.com/Kortexio/ContextMemory.git && cd ContextMemory docker compose -f docker-compose.yml -f docker-compose.llamacpp.yml up --build -d CPU; GPU: docker-compose.vllm.yml ./scripts/aha-chat.sh Windows: .\scripts\aha-chat.ps1 aha-chat sends two requests in the same session. The second one carries only the new question: js == Turn 1: "Remember this for later: our staging database host is postgres-staging-01." == Turn 2: "What is our staging database host?" no history in the request body AHA OK — the client sent no history; the gateway remembered 'postgres-staging-01'. The same check runs in CI against a real llama-server on every push e2e-llamacpp https://github.com/Kortexio/ContextMemory/blob/main/.github/workflows/e2e-llamacpp.yml . First start downloads a ~2 GB GGUF; pick another model with LLAMACPP HF MODEL Engines https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md engines . Admin UI: http://localhost:5200 . ContextMemory https://github.com/Kortexio/ContextMemory is the open-source agentic memory gateway behind Kortexio https://kortexio.io . Your app or Cursor/Claude keeps talking to a normal chat API. The gateway: 1. Authenticates the tenant and attaches session wiki + history 2. Runs an agentic tool loop when tools are enabled wiki search, sandbox, MCP, … 3. Applies skills, guardrails, validators , and optional HITL before destructive actions 4. Returns a standard OpenAI-shaped chat.completions response streaming supported Your client OpenAI SDK / Cursor MCP / curl │ ▼ POST /v1/chat/completions ┌────────────────────────────────────────────┐ │ ContextMemory .NET 9 │ │ Auth · session wiki · Global Wiki tool │ │ Agentic loop · skills · guardrails · HITL │ │ LLM backend per app — BYO engine │ └───────┬──────────────────┬─────────────────┘ ▼ ▼ sandbox-runtime mcp-runtime / MCP servers shell/python/node HTTP + stdio, OAuth or Azure ACA sessions Honest boundaries: this is a gateway + server-side harness , not a client agent framework LangGraph/CrewAI and not an agent OS Letta . You keep your OpenAI client; the loop runs on the server. LLM engines: ContextMemory does not ship or lock to one inference stack. Per tenant you pick any OpenAI-compatible /v1 host — Ollama, vLLM, LM Studio, ExLlamaSharp https://github.com/Kortexio/ExLlamaSharp , OpenAI, Azure-compatible, LiteLLM, custom. Compose ships llama.cpp and vLLM overrides; swap engines per app in Admin → Config → LLM . How we compare Mem0 / Zep / Letta / why we are not RAG : docs/compare.md https://github.com/Kortexio/ContextMemory/blob/main/docs/compare.md . | You need… | ContextMemory provides… | |---|---| | Memory that survives turns without rewriting your client | Session markdown wiki + history inject; send only the new message | | Memory you can open, edit, audit | Files on disk / Postgres — not opaque embeddings | | Shared company/docs knowledge in chat | Global Wiki digests + on-demand wiki search / wiki grep not classic RAG / embeddings | | Tools without a second orchestrator | Same /v1 : sandbox + MCP + wiki tools | | Safer agents | Skills & guardrail packs, validators, HITL CONFIRM:id | | Cursor / Claude permanent memory fast | MCP wedge: memory save / memory search / memory get | | Any LLM per tenant | BYO /v1 — Ollama, vLLM, LM Studio, ExLlamaSharp, OpenAI, Azure-compatible, custom | | Operate without a test client | Admin + Playground | | Full control / zero ops | Docker self-host · Kortexio Cloud https://kortexio.io cmk live … | Full detail: docs/architecture-and-features.md https://github.com/Kortexio/ContextMemory/blob/main/docs/architecture-and-features.md · Admin: docs/admin-ui.md https://github.com/Kortexio/ContextMemory/blob/main/docs/admin-ui.md · HITL: docs/hitl.md https://github.com/Kortexio/ContextMemory/blob/main/docs/hitl.md . | Area | Highlights | |---|---| | Memory | Session wiki + rolling summary; history budgets; Global Wiki digests/FTS/revisions asOf ; no vector RAG | | Agentic | Server-side tool loop; sandbox; MCP catalog; artifacts; subagents; validators; HITL; egress policy | | Skills | Platform + per-app skills/guardrails skill / always on / requestable | | MCP | Outbound wedge Cursor → CM · inbound catalog CM → your MCP servers | | Ops | Admin UI · File or Postgres · Prometheus /metrics · Compose API + Admin + mcp-runtime + sandbox | The gateway talks OpenAI-compatible /v1 to the engine and Ollama native /api/chat when you need num ctx . Change the engine anytime in Admin → Config → LLM or PATCH /admin/apps/{id}/config . docker run --rm -p 5100:8080 \ -v contextmemory-data:/app/data \ -e ContextMemory MasterKey=cm master dev key change me \ -e ContextMemory Apps demo-dev ApiKey=cm live dev key change me \ -e ContextMemory Apps demo-dev LlmBackend=openai-compatible \ -e ContextMemory Apps demo-dev LlmModel=local-model \ -e ContextMemory Apps demo-dev LlmEndpoint=http://host.docker.internal:8080 \ --add-host=host.docker.internal:host-gateway \ ghcr.io/kortexio/contextmemory:latest Host-level default for all apps: ContextMemory LlmEndpoint . Ollama on the host works too ContextMemory LlmEndpoint=http://host.docker.internal:11434 . Engine flags that matter llama.cpp --jinja , vLLM tool parser : Engines https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md engines . Full stack API + Admin + MCP + sandbox : docs/self-host.md https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md . git clone https://github.com/Kortexio/ContextMemory.git cd ContextMemory/mcp-server && npm install && node print-mcp-config.mjs Paste into Cursor → Settings → MCP or ~/.cursor/mcp.json . Same snippet works for Claude Desktop. Details: mcp-server/README.md https://github.com/Kortexio/ContextMemory/blob/main/mcp-server/README.md . Then, in two separate chats: | Chat | You say | Agent should | |---|---|---| | A | Remember: staging DB is postgres-staging-01 | memory save | | B new | What is our staging DB? | memory search + answer | Same flow without Cursor wiki API, no LLM : ./scripts/aha-demo.sh or .\scripts\aha-demo.ps1 . | | Kortexio Cloud https://kortexio.io | Self-host this repo | |---|---|---| | Best for | Zero ops | Full control API + Admin + MCP + sandbox | | Key | cmk live … no X-App-Id | cm live … + X-App-Id | | Chat body | Identical OpenAI /v1 | Identical OpenAI /v1 | | LLM | BYO provider in dashboard | BYO engine in Admin / env | curl -X POST http://localhost:5100/v1/chat/completions \ -H "Content-Type: application/json" \ -H "X-App-Id: demo-dev" -H "X-User-Id: user-42" -H "X-Session-Id: sess-abc" \ -H "Authorization: Bearer cm live dev key change me" \ -d '{"model":"local-model","messages": {"role":"user","content":"Hello"} }' Thin header helpers not full SDKs : @kortexio/contextmemory https://www.npmjs.com/package/@kortexio/contextmemory · kortexio-contextmemory | Doc | Topic | |---|---| | docs/compare.md https://github.com/Kortexio/ContextMemory/blob/main/docs/compare.md | Why it exists · vs Mem0 / Zep / Letta · why we are not RAG | | docs/architecture-and-features.md https://github.com/Kortexio/ContextMemory/blob/main/docs/architecture-and-features.md | Wiki, temporal facts, agentic, skills, LLM backends | | docs/admin-ui.md https://github.com/Kortexio/ContextMemory/blob/main/docs/admin-ui.md | Admin UI map | | docs/hitl.md https://github.com/Kortexio/ContextMemory/blob/main/docs/hitl.md | Human-in-the-loop | | docs/api.md https://github.com/Kortexio/ContextMemory/blob/main/docs/api.md | HTTP API | | docs/cloud.md https://github.com/Kortexio/ContextMemory/blob/main/docs/cloud.md · docs/self-host.md https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md | Cloud · Docker / Compose | | docs/ops.md https://github.com/Kortexio/ContextMemory/blob/main/docs/ops.md | Ops & troubleshooting | | docs/README.md https://github.com/Kortexio/ContextMemory/blob/main/docs/README.md | Full docs index | Website: kortexio.io https://kortexio.io · Email: hello@kortexio.io mailto:hello@kortexio.io AGPL-3.0 for this open-source core — self-host it freely, including commercially. Need to embed it in a closed-source product without AGPL obligations? Use Kortexio Cloud https://kortexio.io or a commercial license. See docs/license-and-support.md https://github.com/Kortexio/ContextMemory/blob/main/docs/license-and-support.md .