{"slug": "show-hn-try-free-long-term-memory-for-ai-50m-token-window", "title": "Show HN: Try Free Long-Term Memory for AI 50M-Token Window", "summary": "Corbenic AI released Galahad, a memory layer for AI models, as a free PyPI package (`pip install galahad-kv`) that stores a model's KV cache to disk so previously processed tokens are not recomputed. The company reports 99.6% of tokens returned from memory, 14× faster inference on vLLM at 0.59 s per question, byte-exact document retrieval scoring 100/100 versus 77 for RAGFlow, and a 50-million-token window with GPU memory flat at 34.1 GB. Galahad runs inside llama.cpp, vLLM and SGLang via a C++ core, is free for 1 GPU under a non-commercial license, and supports Linux x86-64 with Python 3.10–3.14.", "body_md": "**Now live: `pip install galahad-kv`**\n\nGalahad is the memory layer for AI. A model reads a text once, and Galahad keeps that reading. When the text is needed again, Galahad gives the reading back, so the GPU never reads the same text twice. Galahad also keeps the documents themselves and finds the part a question is about, so the model reads only what matters.\n\nGalahad has a C++ core and runs inside llama.cpp, vLLM and SGLang.\n\n- **Don't pay for the same tokens twice.** The tokens a model already processed are\nnot processed again. Galahad gives the reading back instead of recomputing it.**99.6%** of tokens came back from memory;**14× faster** on vLLM,**0.59 s** per question.\n- **Byte-exact, not approximate.** What comes back is identical to what went in: the\nmodel's exact reading, and your documents stored as exact text. No lossy embedding,\nno \"close enough\" like a RAG pipeline.**100/100** right vs**77** for RAGFlow.\n- **Reads only what matters.** Blaise finds the chapter a question is about, so the\nmodel reads**670** tokens instead of**9,700** .\n- **Read past the context window.** A text far larger than the model's context, up to**50 million tokens** measured, is read in parts; each part's reading is saved, and\nthe part a question needs is given back. GPU memory stays**flat at 34.1 GB** whether\nthe text is 1 million or 50 million tokens.\n- **Memory that survives restarts** , encrypted at rest, and your key never leaves your\nmachine.\n\n**Free for 1 GPU.**\n\n```\npip install galahad-kv\n```\n\nThen activate it (one command for everything):\n\n```\ngalahad free --org <your-secret-key> --email you@company.com --accept-non-commercial\ngalahad doctor          # library loads? licence valid? store writable?\n```\n\n- **vLLM:**`vllm serve <model> --kv-transfer-config '{\"kv_connector\":\"GalahadConnector\",\"kv_role\":\"kv_both\"}'`\n- **SGLang:** add`--hicache-storage-backend dynamic` with the Galahad backend (see the[wiki](https://github.com/corbenicai/galahad/wiki/Install) )\n- **Bare / your own code:** link`libgalahad.so` (ships in the wheel)\n\nLinux x86-64, Python 3.10–3.14. On [PyPI](https://pypi.org/project/galahad-kv/).\nFull guide: the [wiki](https://github.com/corbenicai/galahad/wiki).\n\n|  | Feature | What it does | Measured | \n|---|---|---|---|\n| 💾 | **Taliesin** · the memory | Keeps what the model has read. The GPU never reads the same text twice. | **99.6%** of tokens came back from memory | \n| 📚 | **Blaise** · the library | Keeps your documents and finds the chapter a question is about. | **100/100** right, 670 tokens read instead of 9,700 | \n| 🪟 | **50M-token window** · read past the context limit | A text far past the model's context is read in parts; the window moves over disk, the GPU holds one part. | GPU memory flat **34.1 GB** at 1M and at 50M tokens | \n| ⚙️ | **Inside your own program** | The model and Galahad in one C++ program. No server, no Python. | **100/100** right, 0.65 s per question | \n\n## **How it works**\n\n- **Taliesin** saves the model's KV cache to disk and restores it. It plugs in as the vLLM KV connector, the SGLang HiCache storage backend, and llama.cpp slot save and restore.\n- **Blaise** keeps your documents as exact text and returns the chapter a question is about. Text is stored byte-exact.\n- **50M-token window:** a long text is read in parts; each part's reading is saved to disk and the part is dropped from the GPU, so GPU memory stays flat while the window moves over the whole text. The part a question needs is given back and answered from. Nothing to switch on. See the[wiki](https://github.com/corbenicai/galahad/wiki/Performance#very-long-texts) .\n- **Inside your own program:** llama.cpp is built into the Galahad library; your program calls its C API directly.\n\n|  | Feature | What it does | Measured | \n|---|---|---|---|\n| 🔗 | **Agent connectors** | Records every step of your agent. Never changes what it does. | **0** answers changed, no measurable slowdown | \n| 🔁 | **Replay** | Compares a run with a recording and shows the exact step that changed. | One changed word found at **step 0, word 0** | \n| 🎯 | **Fault finder** | Says what caused a failure: your code, the model, or a saved reading. | Right cause on all 3 runtimes in **0.001 ms** | \n| 🌿 | **Branching** | Many tries from one start, without copying it. | **0.18 ms** instead of 269 ms per branch | \n| 🕰️ | **Inspector and rewind** | See what is in memory. Put an agent back to an earlier point. | Rewind passed **27/27** checks | \n\n## **How it works**\n\n- **Agent connectors:** one decorator per step, for LangGraph, LangChain, CrewAI or your own loop. Each run becomes a causal graph.\n- **Replay:** a pytest plugin. Record once; every CI run compares with the recording, so a prompt change that breaks the agent fails the build.\n- **Fault finder:** returns the layer, the rule that decided it and the suspect block. With too little evidence it says \"unknown\" instead of guessing.\n- **Branching:** copy-on-write branches of the KV cache (fork, snapshot, restore, discard). On vLLM and SGLang through a per-request`galahad_branch` flag; snapshots survive a restart.\n- **Inspector and rewind:** a local web page and API behind a token, with an audit log. Rewind must be switched on by the host.\n\n|  | Feature | What it does | Measured | \n|---|---|---|---|\n| 🧩 | **Prefix sharing** | A start that many requests share is read once. | **22×** less to compute for 50 agents | \n| 📌 | **Pinning** | Keeps what every request needs, such as a system prompt. | Survived **181** evictions | \n| 🏢 | **Tenants** | One server, many customers, each kept apart and counted. | Identical questions from two customers stayed apart | \n\n## **How it works**\n\n- **Prefix sharing:** the shared start is saved as a chain of chunks, and a lookup finds the longest saved start. It is used only when loading is faster than computing.\n- **Pinning:** pinned blocks are skipped when Galahad has to free space.\n- **Tenants:** each tenant has its own keyspace, with hits, misses and evictions counted per tenant.\n\n|  | Feature | What it does | Measured | \n|---|---|---|---|\n| 🔐 | **Encryption at rest** | Everything saved is encrypted. A forgotten customer is gone for good. | **0 bytes** readable on disk | \n| 🛡️ | **Sanitizer** | Refuses a broken reading before it is saved. | Checked in **0.05 ms** | \n| 💰 | **FinOps** | Shows what the memory saved, in GPU time and money. It never guesses. | **746–960** GPU-seconds saved per 100 questions | \n\n## **How it works**\n\n- **Encryption at rest:** AES-256-GCM on every saved block and on the catalogue, with keys from your own key service. With a wrong key, the store refuses to open.\n- **Sanitizer:** checks for NaN and infinity in the number format the host declares (bf16, f16), and checks the size.\n- **FinOps:** uses your GPU price per hour and usage level. Without a price it gives no number and says what is missing.\n\n<sub>Measured with Gemma 4 31B on llama.cpp, vLLM and SGLang, September 2026.</sub>\n\nGalahad is tested on **30 open models from 10 families**, from 7B to 70B.\n\n| Family | Models | \n|---|---|\n| **Qwen** | Qwen3 8B · Qwen3 14B · Qwen3 32B · Qwen3.5 9B · Qwen3.8 27B · Qwen3.6 35B-A3B · Qwen3 Coder 30B-A3B | \n| **Gemma** | Gemma 4 12B · Gemma 4 26B-A4B · Gemma 4 31B | \n| **Mistral** | Ministral 3 8B · Ministral 3 14B · Mistral Small 3.2 24B · Devstral Small 2 24B | \n| **GLM** | GLM-4.6V-Flash 9B · GLM-4.7-Flash 30B-A3B | \n| **Llama** | Llama 3.1 8B Instruct · Llama 3.3 70B Instruct | \n| **DeepSeek** | R1 Distill Qwen 7B · R1 Distill Qwen 14B · R1 Distill Qwen 32B · R1 Distill Llama 8B · R1 Distill Llama 70B · R1-0528-Qwen3-8B | \n| **Phi** | Phi-4 14B · Phi-4 Reasoning 14B | \n| **GPT-oss** | GPT-oss 20B | \n| **Nemotron** | Nemotron 3 Nano 30B-A3B | \n| **Olmo** | Olmo 3 7B · Olmo 3 32B | \n\n<sub>© 2026 Corbenic AI - Sietse Schelpe. Patent pending.</sub>", "url": "https://wpnews.pro/news/show-hn-try-free-long-term-memory-for-ai-50m-token-window", "canonical_source": "https://github.com/corbenicai/galahad", "published_at": "2026-10-10 21:36:50+00:00", "updated_at": "2026-10-10 21:47:30.552711+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "mlops", "ai-agents"], "entities": ["Galahad", "Corbenic AI", "llama.cpp", "vLLM", "SGLang", "Taliesin", "Blaise", "RAGFlow"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-try-free-long-term-memory-for-ai-50m-token-window", "markdown": "https://wpnews.pro/news/show-hn-try-free-long-term-memory-for-ai-50m-token-window.md", "text": "https://wpnews.pro/news/show-hn-try-free-long-term-memory-for-ai-50m-token-window.txt", "jsonld": "https://wpnews.pro/news/show-hn-try-free-long-term-memory-for-ai-50m-token-window.jsonld"}}