Show HN: Try Free Long-Term Memory for AI 50M-Token Window Corbenic AI released Galahad, a memory layer for AI models, as a free PyPI package (`pip install galahad-kv`) that stores a model's KV cache to disk so previously processed tokens are not recomputed. The company reports 99.6% of tokens returned from memory, 14× faster inference on vLLM at 0.59 s per question, byte-exact document retrieval scoring 100/100 versus 77 for RAGFlow, and a 50-million-token window with GPU memory flat at 34.1 GB. Galahad runs inside llama.cpp, vLLM and SGLang via a C++ core, is free for 1 GPU under a non-commercial license, and supports Linux x86-64 with Python 3.10–3.14. Now live: pip install galahad-kv Galahad is the memory layer for AI. A model reads a text once, and Galahad keeps that reading. When the text is needed again, Galahad gives the reading back, so the GPU never reads the same text twice. Galahad also keeps the documents themselves and finds the part a question is about, so the model reads only what matters. Galahad has a C++ core and runs inside llama.cpp, vLLM and SGLang. - Don't pay for the same tokens twice. The tokens a model already processed are not processed again. Galahad gives the reading back instead of recomputing it. 99.6% of tokens came back from memory; 14× faster on vLLM, 0.59 s per question. - Byte-exact, not approximate. What comes back is identical to what went in: the model's exact reading, and your documents stored as exact text. No lossy embedding, no "close enough" like a RAG pipeline. 100/100 right vs 77 for RAGFlow. - Reads only what matters. Blaise finds the chapter a question is about, so the model reads 670 tokens instead of 9,700 . - Read past the context window. A text far larger than the model's context, up to 50 million tokens measured, is read in parts; each part's reading is saved, and the part a question needs is given back. GPU memory stays flat at 34.1 GB whether the text is 1 million or 50 million tokens. - Memory that survives restarts , encrypted at rest, and your key never leaves your machine. Free for 1 GPU. pip install galahad-kv Then activate it one command for everything : galahad free --org