Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute A memory layer called galahad-kv saves the key-value (KV) state of each roughly 16,000-token block to encrypted local NVMe disk and reloads it byte-exact without recomputation, and was tested on 50,000,000 tokens. The public package targets the fact that large language models can only use text fitting in their context window and recompute KV state for a prompt on every send. The test claims the approach is faster and cheaper than recompute for long-term memory. Reddit https://www.reddit.com/r/huggingface/s/x2OAv4Xgee . A large language model can only use the text that fits in its context window, and it recomputes its internal key-value KV state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens