cd /news/ai-agents/contextmemory-markdown-memory-for-yo… · home › topics › ai-agents › article
[ARTICLE · art-141561] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

ContextMemory – Markdown memory for your llama.cpp/vLLM server

Kortexio released ContextMemory, an open-source self-hosted memory gateway that sits in front of any OpenAI-compatible /v1 engine such as llama.cpp, vLLM, Ollama, LM Studio or OpenAI and stores session memory as editable markdown files rather than a vector database. The gateway authenticates tenants, injects a session wiki plus history, runs an agentic tool loop with sandbox and MCP tools, applies skills, guardrails, validators and optional human-in-the-loop confirmation, and returns a standard OpenAI-shaped chat.completions response with streaming support. It is installed via git clone and docker compose with llama.cpp or vLLM overrides, and its CI runs an end-to-end check against a real llama-server on every push.

read5 min views1 publishedSep 29, 2026
ContextMemory – Markdown memory for your llama.cpp/vLLM server
Image: Michielbdejong (auto-discovered)

Try it (3 commands) · Engines · Docs · vs Mem0 / Zep / Letta

Self-hosted memory gateway for your llama.cpp / vLLM server.

Put one OpenAI-compatible /v1 URL in front of your engine. Your client sends only the new message; the gateway keeps session memory as markdown you can open, edit, and diff. No vector DB, no client rewrite.

git clone https://github.com/Kortexio/ContextMemory.git && cd ContextMemory
docker compose -f docker-compose.yml -f docker-compose.llamacpp.yml up --build -d   # CPU; GPU: docker-compose.vllm.yml
./scripts/aha-chat.sh                                                             # Windows: .\scripts\aha-chat.ps1

aha-chat sends two requests in the same session. The second one carries only the new question:

==> Turn 1: "Remember this for later: our staging database host is postgres-staging-01."
==> Turn 2: "What is our staging database host?"   (no history in the request body)
AHA OK — the client sent no history; the gateway remembered 'postgres-staging-01'.

The same check runs in CI against a real llama-server on every push (e2e-llamacpp). First start downloads a ~2 GB GGUF; pick another model with LLAMACPP_HF_MODEL (Engines). Admin UI: http://localhost:5200.

ContextMemory is the open-source agentic memory gateway behind Kortexio.

Your app (or Cursor/Claude) keeps talking to a normal chat API. The gateway:

  1. Authenticates the tenant and attaches session wiki + history
  2. Runs an agentic tool loop when tools are enabled (wiki search, sandbox, MCP, …)
  3. Applies skills, guardrails, validators , and optionalHITL before destructive actions
  4. Returns a standard OpenAI-shaped chat.completions response (streaming supported)
Your client (OpenAI SDK / Cursor MCP / curl)
        │
        ▼  POST /v1/chat/completions
┌────────────────────────────────────────────┐
│  ContextMemory (.NET 9)                    │
│  Auth · session wiki · Global Wiki tool    │
│  Agentic loop · skills · guardrails · HITL │
│  LLM backend (per app — BYO engine)        │
└───────┬──────────────────┬─────────────────┘
        ▼                  ▼
 sandbox-runtime      mcp-runtime / MCP servers
 (shell/python/node)  (HTTP + stdio, OAuth)
 or Azure ACA sessions

Honest boundaries: this is a gateway + server-side harness, not a client agent framework (LangGraph/CrewAI) and not an agent OS (Letta). You keep your OpenAI client; the loop runs on the server.

LLM engines: ContextMemory does not ship or lock to one inference stack. Per tenant you pick any OpenAI-compatible /v1 host — Ollama, vLLM, LM Studio, ExLlamaSharp, OpenAI, Azure-compatible, LiteLLM, custom. Compose ships llama.cpp and vLLM overrides; swap engines per app in Admin → Config → LLM.

How we compare (Mem0 / Zep / Letta / why we are not RAG): docs/compare.md.

You need… ContextMemory provides…
Memory that survives turns without rewriting your client Session markdown wiki + history inject; send only the new message
Memory you can open, edit, audit Files on disk / Postgres — not opaque embeddings
Shared company/docs knowledge in chat Global Wiki digests + on-demandwiki_search /wiki_grep (not classic RAG / embeddings)
Tools without a second orchestrator Same /v1 : sandbox + MCP + wiki tools
Safer agents Skills & guardrail packs, validators, HITL [CONFIRM:id]
Cursor / Claude permanent memory fast MCP wedge: memory_save /memory_search /memory_get
Any LLM per tenant BYO /v1 — Ollama, vLLM, LM Studio, ExLlamaSharp, OpenAI, Azure-compatible, custom
Operate without a test client Admin +Playground
Full control / zero ops Docker self-host · Kortexio Cloud (cmk_live_… )

Full detail: docs/architecture-and-features.md · Admin: docs/admin-ui.md · HITL: docs/hitl.md.

Area Highlights
Memory Session wiki + rolling summary; history budgets; Global Wiki digests/FTS/revisions ( asOf );no vector RAG
Agentic Server-side tool loop; sandbox; MCP catalog; artifacts; subagents; validators; HITL; egress policy
Skills Platform + per-app skills/guardrails ( skill /always_on /requestable )
MCP Outbound wedge (Cursor → CM) · inbound catalog (CM → your MCP servers)
Ops Admin UI · File or Postgres · Prometheus /metrics · Compose (API + Admin + mcp-runtime + sandbox)

The gateway talks OpenAI-compatible /v1 to the engine (and Ollama native /api/chat when you need num_ctx). Change the engine anytime in Admin → Config → LLM or PATCH /admin/apps/{id}/config.

docker run --rm -p 5100:8080 \
  -v contextmemory-data:/app/data \
  -e ContextMemory__MasterKey=cm_master_dev_key_change_me \
  -e ContextMemory__Apps__demo-dev__ApiKey=cm_live_dev_key_change_me \
  -e ContextMemory__Apps__demo-dev__LlmBackend=openai-compatible \
  -e ContextMemory__Apps__demo-dev__LlmModel=local-model \
  -e ContextMemory__Apps__demo-dev__LlmEndpoint=http://host.docker.internal:8080 \
  --add-host=host.docker.internal:host-gateway \
  ghcr.io/kortexio/contextmemory:latest

Host-level default for all apps: ContextMemory__LlmEndpoint. Ollama on the host works too ( ContextMemory__LlmEndpoint=http://host.docker.internal:11434). Engine flags that matter (llama.cpp --jinja, vLLM tool parser): Engines. Full stack (API + Admin + MCP + sandbox): docs/self-host.md.

git clone https://github.com/Kortexio/ContextMemory.git
cd ContextMemory/mcp-server && npm install && node print-mcp-config.mjs

Paste into Cursor → Settings → MCP (or ~/.cursor/mcp.json). Same snippet works for Claude Desktop. Details: mcp-server/README.md.

Then, in two separate chats:

Chat You say Agent should
A Remember: staging DB is postgres-staging-01 memory_save
B (new) What is our staging DB? memory_search + answer

Same flow without Cursor (wiki API, no LLM): ./scripts/aha-demo.sh or .\scripts\aha-demo.ps1.

Kortexio Cloud Self-host (this repo)
Best for Zero ops Full control (API + Admin + MCP + sandbox)
Key cmk_live_… (noX-App-Id ) cm_live_… +X-App-Id
Chat body Identical OpenAI /v1 Identical OpenAI /v1
LLM BYO provider in dashboard BYO engine in Admin / env
curl -X POST http://localhost:5100/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "X-App-Id: demo-dev" -H "X-User-Id: user-42" -H "X-Session-Id: sess-abc" \
  -H "Authorization: Bearer cm_live_dev_key_change_me" \
  -d '{"model":"local-model","messages":[{"role":"user","content":"Hello"}]}'

Thin header helpers (not full SDKs): @kortexio/contextmemory · kortexio-contextmemory

Doc Topic
docs/compare.md Why it exists · vs Mem0 / Zep / Letta · why we are not RAG
docs/architecture-and-features.md Wiki, temporal facts, agentic, skills, LLM backends
docs/admin-ui.md Admin UI map
docs/hitl.md Human-in-the-loop
docs/api.md HTTP API
docs/cloud.md ·docs/self-host.md Cloud · Docker / Compose
docs/ops.md Ops & troubleshooting
docs/README.md Full docs index

Website: kortexio.io · Email: hello@kortexio.io

AGPL-3.0 for this open-source core — self-host it freely, including commercially. Need to embed it in a closed-source product without AGPL obligations? Use Kortexio Cloud or a commercial license. See docs/license-and-support.md.

── more in #ai-agents 4 stories · sorted by recency
── more on @kortexio 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/contextmemory-markdo…] indexed:0 read:5min 2026-09-29 · —