OptChat: An endless chat where the AI remembers everything A developer has released OptChat, a chat harness that treats the entire conversation history as persistent agent memory, stored as a compressed binary tree of one-line summaries. Each turn starts from a fresh context containing a fixed-size (~64k token) view of the whole chat, with a "zoom" tool that expands any summary line back down to the original message, eliminating context rot and manual compaction. The design builds on the earlier OptMem memory tool and is published as a full technical specification for others to implement. OptChat: an endless chat where the AI remembers everything What this is AI agents forget. A chat session fills up, gets compacted or thrown away, and the next one starts from zero. If you work with agents every day, your life ends up scattered over hundreds of sessions in different tools, and the agent can't find any of it: not the decision you made last month, not the algorithm you designed in March, not the PR you forgot to answer. Long sessions have a second problem: context rot . The longer the context, the worse the model works. Compaction "summarize and keep going" throws detail away for good, and the result keeps decaying. OptChat fixes both with one idea: the chat history itself is the memory , stored as a compressed tree. - There is ONE chat, and it never ends. Every message yours, the agent's replies, its tool calls and their results is appended to a log and kept forever, word for word. - In the background, a cheap model compresses the log into a binary tree of one-line summaries: each message becomes a line, two adjacent lines merge into one line covering both, two of those merge again, and so on. - Each time you send a message, the agent starts FRESH: no leftover context. It sees a fixed-size "view" about 64k tokens of the WHOLE chat: recent messages one line each, older ones many per line, the older the coarser. Then it sees your message. - When a line is too vague, the agent "zooms": it opens the line into the two lines it was made from, down to the original message. Any fact from your whole history is a few zooms away. What you get: - Infinite context with a constant size. Nothing is ever deleted; only the resolution of the distant past fades. - No context rot and no manual compaction. Every turn starts clean. - Low cost. Most tokens of an agent turn are its own tool loop, and that is prompt-cached. Across turns, the view changes only near its end, so most of it is cached too. - Your instructions stick. Corrections and preferences you give in chat stay in the memory the compressor ranks your own words first , so most of an AGENTS.md becomes unnecessary: change your mind in a message and the latest ruling wins. - You can browse it too. The tree is a plain file you can walk from the root down to any message. OptChat grew out of OptMem github.com/VictorTaelin/OptMem , a memory tool: an append-only log of short notes plus a summary tree, which an agent reads at the start of each session. OptChat turns that around: instead of a tool the agent calls, the memory IS the chat, and the harness builds every turn from it. The rest of this document is a complete technical spec, for anyone human or AI who wants to build it. It describes a working implementation in full, including the reasons behind each choice. Many choices that look natural are wrong they break the cache, or make the memory decay, or make the agent act on partial text , and the spec says which and why. If you are an AI building this: follow it exactly, and when you deviate, have a reason that the spec does not already refute. --- Technical specification 1. Overview Components: 1. The log ROOT : every message, verbatim, append-only. 2. The tree : summaries. Node l, i covers messages i·2^l, i+1 ·2^l . Level 0 summarizes one message; level l 0 merges its two children l-1, 2i and l-1, 2i+1 . 3. The compactor : a background worker that builds tree nodes with a cheap model, in a strict order. 4. The view : a list of tree nodes that covers the whole chat, oldest first, kept under a byte budget, changed incrementally. 5. The turn loop : each user message starts a fresh model call whose input is system prompt view new message , plus tools zoom and date . 6. Caching : the request layout and cache breakpoints that make all of this cheap. Constants used in the reference implementation: | name | value | meaning | |---|---|---| | NODE | 512 bytes | target size of one summary line | | VIEW | 128,000 bytes | budget of the view ≈ 62-64k tokens | | JOBS | 8 | compactor calls running at once | | TRIES | 5 | attempts per node to get under NODE | | RETRY | 10 s | wait before retrying a failed node | | CAP | 30,000 chars | max size of one tool result head + tail kept | | MARKS | 50,000 / 80,000 / 100,000 chars | cache breakpoints inside the view | All sizes are UTF-8 bytes or characters, for the cache marks , never tokens: a tokenizer changes between models, a byte count never does. 2. Storage Two append-only streams in one directory: chat/ main/YYYY-MM-DD.jsonl one message per line: {i, kind, text, size, date} tree/YYYY-MM-DD.jsonl one node per line: {l, i, text, size} - i is the message index 0, 1, 2, ... , its permanent id. kind is one of: user the user's words; also subagent reports, see §9 , talk the agent's replies , tool the agent's tool calls, as text: name and JSON input , echo tool results , note memories imported from an older system . date is ISO time. size = bytes of kind + ": " + text . - A line goes to the file of the local day it was written. Files split by day only to keep them manageable; the ids are global. - Durability : each line is written with one write and then fsync , before the function returns. A crash loses nothing written. - Torn lines : at load, a line that is not valid JSON a crash mid-write is reported and skipped, and a file not ending in \n gets one appended, so the next write starts on its own line. - One writer : two processes on the same chat would corrupt it. Hold a lock for the life of the process. The reference uses a Unix socket: the process listens on lock ; a second process that can connect to it exits; a socket that refuses connections is stale the OS frees it when the owner dies , so it is deleted and taken over. No PID files, no timeouts. - The tree is a cache in principle rebuildable from the log , but it costs model calls to rebuild, so it is stored and never recomputed. Never edit or delete anything in these files. The log is history. Thoughts model reasoning are shown to the user but never logged . Reason: the compactor would have to summarize them, and Claude Sonnet's reasoning-extraction safeguard refused the compactor on thoughts in 112 of 630 first tries; without thoughts, 0 of 523. Thoughts also add little that the replies and tool calls don't already show. 3. The tree node 0, i = message i, in ≤ NODE bytes node l, i = merge node l-1, 2i , node l-1, 2i+1 , in ≤ NODE bytes covers l,i = messages i·2^l, i+1 ·2^l The tree is purely binary . OptMem built small blocks of up to 16 memories straight from the raw notes, and only bigger ones from two halves. That had no real justification and made the structure confusing; OptChat dropped it. Every parent comes from exactly its two children. Free nodes. If the source already fits in NODE bytes, it IS the node, with no model call: - level 0: kind + ": " + text of a short message, verbatim so short user messages stay word for word forever, up to the level where they get merged ; - level 0: childA + "\n" + childB , if that fits. Why 512 bytes. It started at 128 and the model couldn't write useful lines that short and overshot often . OptMem used 280. 512 bytes is a dense paragraph: room for several items with names and numbers. With an average real line around 250 bytes, a 128 KB view holds ~500 lines. Addressing. A node is named id+n : id = its first message, n = 2^l = how many messages it covers. So node l, i is i·2^l + 2^l . Real message ids, not tree coordinates: the agent reads 2184+8 in the view and calls zoom 2184, 8 directly. 4. The compactor 4.1 When a node is built A background loop "pump" scans all levels and starts every node that: 1. is not built and not already running; 2. has its sources: level 0, the message exists; level 0, both children are built; 3. has its whole context summarized: every line of the current view that lies before the node's end is a built summary see the first function below . Up to JOBS run at once. After each finishes or fails , pump again. Rule 3 is essential, and it gives OptMem's order for free: messages are compressed one at a time, in order , while merges of finished parts run alongside. The compactor never sees a line that isn't a summary. function first mem : first message whose view line is unbuilt for part in mem.view: if not built part : return start part return len mem.root function pump mem : T = len mem.root for l = 0 while 2^l <= T: for i = 0 while i+1 ·2^l <= T: if len busy = JOBS: return end = l == 0 ? i : i+1 ·2^l if built l,i or busy l,i or not ready l,i or end first mem : continue busy.add l,i build l,i .then ok - busy.remove l,i ; fail.remove l,i ; pump mem , err - report err once per node; after RETRY: busy.remove l,i ; pump mem For level 0 node i , end = i means "every view line before message i is a summary", so node i is the first unbuilt one. For a merge, end = i+1 ·2^l means everything up to the node's last message is summarized. Retry : a failed node waits RETRY 10 s and is tried again, forever; only its first failure is reported. Don't use long exponential backoff: the next turn waits for these summaries §6 , so the compactor must catch up as fast as possible. 4.2 What a compactor call sees One call per node, with no tools, on a cheap model the reference uses Claude Sonnet at medium effort; at low effort it overshot the size limit much more . Input: - system : the COMPACT prompt §4.4 , constant. - user message , two text blocks: 1. Context : the current view's lines up to the node, wrapped in