{"slug": "contextmemory-markdown-memory-for-your-llama-cpp-vllm-server", "title": "ContextMemory – Markdown memory for your llama.cpp/vLLM server", "summary": "Kortexio released ContextMemory, an open-source self-hosted memory gateway that sits in front of any OpenAI-compatible /v1 engine such as llama.cpp, vLLM, Ollama, LM Studio or OpenAI and stores session memory as editable markdown files rather than a vector database. The gateway authenticates tenants, injects a session wiki plus history, runs an agentic tool loop with sandbox and MCP tools, applies skills, guardrails, validators and optional human-in-the-loop confirmation, and returns a standard OpenAI-shaped chat.completions response with streaming support. It is installed via git clone and docker compose with llama.cpp or vLLM overrides, and its CI runs an end-to-end check against a real llama-server on every push.", "body_md": "[**Try it (3 commands)**](#try-it-in-three-commands)\n  ·\n  [Engines](https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md#engines)\n  ·\n  [Docs](https://github.com/Kortexio/ContextMemory/blob/main/docs/README.md)\n  ·\n  [vs Mem0 / Zep / Letta](https://github.com/Kortexio/ContextMemory/blob/main/docs/compare.md)\n\n**Self-hosted memory gateway for your llama.cpp / vLLM server.**\n\n  Put one OpenAI-compatible `/v1` URL in front of your engine. Your client sends only the new message;\n  the gateway keeps session memory as **markdown you can open, edit, and diff**. No vector DB, no client rewrite.\n\n```\ngit clone https://github.com/Kortexio/ContextMemory.git && cd ContextMemory\ndocker compose -f docker-compose.yml -f docker-compose.llamacpp.yml up --build -d   # CPU; GPU: docker-compose.vllm.yml\n./scripts/aha-chat.sh                                                             # Windows: .\\scripts\\aha-chat.ps1\n```\n\n`aha-chat` sends two requests in the same session. The second one carries **only** the new question:\n\n``` js\n==> Turn 1: \"Remember this for later: our staging database host is postgres-staging-01.\"\n==> Turn 2: \"What is our staging database host?\"   (no history in the request body)\nAHA OK — the client sent no history; the gateway remembered 'postgres-staging-01'.\n```\n\nThe same check runs in CI against a real `llama-server` on every push ([`e2e-llamacpp`](https://github.com/Kortexio/ContextMemory/blob/main/.github/workflows/e2e-llamacpp.yml)). First start downloads a ~2 GB GGUF; pick another model with `LLAMACPP_HF_MODEL` ([Engines](https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md#engines)). Admin UI: `http://localhost:5200`.\n\n[ContextMemory](https://github.com/Kortexio/ContextMemory) is the open-source **agentic memory gateway** behind [Kortexio](https://kortexio.io).\n\nYour app (or Cursor/Claude) keeps talking to a normal chat API. The gateway:\n\n1. Authenticates the tenant and attaches **session wiki + history**\n2. Runs an **agentic tool loop** when tools are enabled (wiki search, sandbox, MCP, …)\n3. Applies **skills, guardrails, validators** , and optional**HITL** before destructive actions\n4. Returns a standard OpenAI-shaped `chat.completions` response (streaming supported)\n\n```\nYour client (OpenAI SDK / Cursor MCP / curl)\n        │\n        ▼  POST /v1/chat/completions\n┌────────────────────────────────────────────┐\n│  ContextMemory (.NET 9)                    │\n│  Auth · session wiki · Global Wiki tool    │\n│  Agentic loop · skills · guardrails · HITL │\n│  LLM backend (per app — BYO engine)        │\n└───────┬──────────────────┬─────────────────┘\n        ▼                  ▼\n sandbox-runtime      mcp-runtime / MCP servers\n (shell/python/node)  (HTTP + stdio, OAuth)\n or Azure ACA sessions\n```\n\n**Honest boundaries:** this is a **gateway + server-side harness**, not a client agent framework (LangGraph/CrewAI) and not an agent OS (Letta). You keep your OpenAI client; the loop runs on the server.\n\n**LLM engines:** ContextMemory does **not** ship or lock to one inference stack. Per tenant you pick any OpenAI-compatible `/v1` host — Ollama, vLLM, LM Studio, [ExLlamaSharp](https://github.com/Kortexio/ExLlamaSharp), OpenAI, Azure-compatible, LiteLLM, custom. Compose ships llama.cpp and vLLM overrides; swap engines per app in **Admin → Config → LLM**.\n\nHow we compare (Mem0 / Zep / Letta / **why we are not RAG**): [`docs/compare.md`](https://github.com/Kortexio/ContextMemory/blob/main/docs/compare.md).\n\n| You need… | ContextMemory provides… | \n|---|---|\n| Memory that survives turns without rewriting your client | Session markdown wiki + history inject; send only the new message | \n| Memory you can open, edit, audit | Files on disk / Postgres — not opaque embeddings | \n| Shared company/docs knowledge in chat | **Global Wiki** digests + on-demand`wiki_search` /`wiki_grep` (**not** classic RAG / embeddings) | \n| Tools without a second orchestrator | Same `/v1` : sandbox + MCP + wiki tools | \n| Safer agents | Skills & guardrail packs, validators, HITL `[CONFIRM:id]` | \n| Cursor / Claude permanent memory fast | MCP wedge: `memory_save` /`memory_search` /`memory_get` | \n| Any LLM per tenant | BYO `/v1` — Ollama, vLLM, LM Studio, ExLlamaSharp, OpenAI, Azure-compatible, custom | \n| Operate without a test client | **Admin** +**Playground** | \n| Full control / zero ops | Docker self-host · [Kortexio Cloud](https://kortexio.io) (`cmk_live_…` ) | \n\nFull detail: [`docs/architecture-and-features.md`](https://github.com/Kortexio/ContextMemory/blob/main/docs/architecture-and-features.md) · Admin: [`docs/admin-ui.md`](https://github.com/Kortexio/ContextMemory/blob/main/docs/admin-ui.md) · HITL: [`docs/hitl.md`](https://github.com/Kortexio/ContextMemory/blob/main/docs/hitl.md).\n\n| Area | Highlights | \n|---|---|\n| **Memory** | Session wiki + rolling summary; history budgets; Global Wiki digests/FTS/revisions ( `asOf` );**no vector RAG** | \n| **Agentic** | Server-side tool loop; sandbox; MCP catalog; artifacts; subagents; validators; HITL; egress policy | \n| **Skills** | Platform + per-app skills/guardrails ( `skill` /`always_on` /`requestable` ) | \n| **MCP** | Outbound wedge (Cursor → CM) · inbound catalog (CM → your MCP servers) | \n| **Ops** | Admin UI · File or Postgres · Prometheus `/metrics` · Compose (API + Admin + mcp-runtime + sandbox) | \n\nThe gateway talks OpenAI-compatible `/v1` to the engine (and Ollama native `/api/chat` when you need `num_ctx`). Change the engine anytime in **Admin → Config → LLM** or `PATCH /admin/apps/{id}/config`.\n\n```\ndocker run --rm -p 5100:8080 \\\n  -v contextmemory-data:/app/data \\\n  -e ContextMemory__MasterKey=cm_master_dev_key_change_me \\\n  -e ContextMemory__Apps__demo-dev__ApiKey=cm_live_dev_key_change_me \\\n  -e ContextMemory__Apps__demo-dev__LlmBackend=openai-compatible \\\n  -e ContextMemory__Apps__demo-dev__LlmModel=local-model \\\n  -e ContextMemory__Apps__demo-dev__LlmEndpoint=http://host.docker.internal:8080 \\\n  --add-host=host.docker.internal:host-gateway \\\n  ghcr.io/kortexio/contextmemory:latest\n```\n\nHost-level default for all apps: `ContextMemory__LlmEndpoint`. Ollama on the host works too (` ContextMemory__LlmEndpoint=http://host.docker.internal:11434`). Engine flags that matter (llama.cpp `--jinja`, vLLM tool parser): [Engines](https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md#engines). Full stack (API + Admin + MCP + sandbox): [`docs/self-host.md`](https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md).\n\n```\ngit clone https://github.com/Kortexio/ContextMemory.git\ncd ContextMemory/mcp-server && npm install && node print-mcp-config.mjs\n```\n\nPaste into **Cursor → Settings → MCP** (or `~/.cursor/mcp.json`). Same snippet works for Claude Desktop. Details: [` mcp-server/README.md`](https://github.com/Kortexio/ContextMemory/blob/main/mcp-server/README.md).\n\nThen, in two separate chats:\n\n| Chat | You say | Agent should | \n|---|---|---|\n| **A** | `Remember: staging DB is postgres-staging-01` | `memory_save` | \n| **B** (new) | `What is our staging DB?` | `memory_search` + answer | \n\nSame flow without Cursor (wiki API, no LLM): `./scripts/aha-demo.sh` or `.\\scripts\\aha-demo.ps1`.\n\n|  | **[Kortexio Cloud](https://kortexio.io)** | **Self-host (this repo)** | \n|---|---|---|\n| Best for | Zero ops | Full control (API + Admin + MCP + sandbox) | \n| Key | `cmk_live_…` (no`X-App-Id` ) | `cm_live_…` +`X-App-Id` | \n| Chat body | Identical OpenAI `/v1` | Identical OpenAI `/v1` | \n| LLM | BYO provider in dashboard | BYO engine in Admin / env | \n\n```\ncurl -X POST http://localhost:5100/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -H \"X-App-Id: demo-dev\" -H \"X-User-Id: user-42\" -H \"X-Session-Id: sess-abc\" \\\n  -H \"Authorization: Bearer cm_live_dev_key_change_me\" \\\n  -d '{\"model\":\"local-model\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nThin header helpers (**not** full SDKs): [`@kortexio/contextmemory`](https://www.npmjs.com/package/@kortexio/contextmemory) · `kortexio-contextmemory`\n\n| Doc | Topic | \n|---|---|\n| [docs/compare.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/compare.md) | Why it exists · vs Mem0 / Zep / Letta · **why we are not RAG** | \n| [docs/architecture-and-features.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/architecture-and-features.md) | Wiki, temporal facts, agentic, skills, **LLM backends** | \n| [docs/admin-ui.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/admin-ui.md) | Admin UI map | \n| [docs/hitl.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/hitl.md) | Human-in-the-loop | \n| [docs/api.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/api.md) | HTTP API | \n| [docs/cloud.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/cloud.md) ·[docs/self-host.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/self-host.md) | Cloud · Docker / Compose | \n| [docs/ops.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/ops.md) | Ops & troubleshooting | \n| [docs/README.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/README.md) | Full docs index | \n\nWebsite: [kortexio.io](https://kortexio.io) · Email: [hello@kortexio.io](mailto:hello@kortexio.io)\n\n**AGPL-3.0** for this open-source core — self-host it freely, including commercially. Need to embed it in a closed-source product without AGPL obligations? Use [Kortexio Cloud](https://kortexio.io) or a commercial license. See [docs/license-and-support.md](https://github.com/Kortexio/ContextMemory/blob/main/docs/license-and-support.md).", "url": "https://wpnews.pro/news/contextmemory-markdown-memory-for-your-llama-cpp-vllm-server", "canonical_source": "https://github.com/Kortexio/ContextMemory", "published_at": "2026-09-29 07:25:09+00:00", "updated_at": "2026-09-29 07:48:02.103743+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "large-language-models", "ai-tools", "agent-protocols"], "entities": ["Kortexio", "ContextMemory", "llama.cpp", "vLLM", "Ollama", "LM Studio", "Mem0", "Letta"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/contextmemory-markdown-memory-for-your-llama-cpp-vllm-server", "markdown": "https://wpnews.pro/news/contextmemory-markdown-memory-for-your-llama-cpp-vllm-server.md", "text": "https://wpnews.pro/news/contextmemory-markdown-memory-for-your-llama-cpp-vllm-server.txt", "jsonld": "https://wpnews.pro/news/contextmemory-markdown-memory-for-your-llama-cpp-vllm-server.jsonld"}}