{"slug": "long-term-memory-for-ai-coding-agents-as-a-tree-of-plain-files", "title": "Long-term memory for AI coding agents as a tree of plain files", "summary": "IHMT-MEMORY, a new long-term memory system for AI coding agents, stores decisions, preferences and corrections as a tree of plain text files on the user's disk and retrieves typically 200–900 tokens per query whether the memory holds 50 entries or 50,000, according to its project documentation. The tool is officially supported on Claude Code and tested with Codex and opencode, runs as a standard stdio MCP server, and requires Python 3.10 or newer plus git, with macOS and Linux (Ubuntu) tested and Windows instructions included but untested. Memory is shared across agents and models from Anthropic, OpenAI, Google and Meta, keeps superseded facts flagged OUTDATED as history, and sends nothing to the cloud.", "body_md": "**Long-term memory for your AI coding agents.** Tell your agent something once — a decision, how\nyour setup works, a correction — and it remembers it in every future session, in any project, with\nany of your agents.\n\n- **Remembers across sessions.** Decisions and their reasons, your environment, your preferences,\npeople and projects, corrections. Your agent searches the memory before answering and saves what\nlasts, so you stop repeating yourself.\n- **One memory for all your agents and models.** Claude Code, Codex and opencode can share the same\nmemory: what one saves, the others find. Tested with models from Anthropic, OpenAI, Google and Meta.\n- **Saves tokens.** Instead of pasting your notes or re-explaining context every session, the agent\nretrieves only what the question needs — typically 200–900 tokens, whether the memory holds 50\nentries or 50,000, because search walks a tree instead of reading everything.[Honest numbers](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#11-honest-numbers) , including where it does*not* save.\n- **Understands time.** When something changes (\"I moved to Valencia\", \"staging is on PostgreSQL 17\nnow\"), the old memory is kept as history, flagged`OUTDATED` , and searches answer with the current\none first. Where you live, where you work and your stack are tracked automatically; any other\ncorrection is linked when the agent saves it with`replaces` . When a question is ambiguous (\"Luis\"\n— which one?), it asks instead of guessing.\n- **Portable.** Your memory is one folder of plain text files. Copy it to another computer, back it\nup, or put it under version control — it works wherever you put it.\n- **Local, private and readable.** No cloud, no database, no account: IHMT stores everything on your\ndisk and sends nothing anywhere. (The memories your agent retrieves reach its model like any other\ncontext.) Every memory is a text file you can open, and each person who installs IHMT starts with\ntheir own, empty memory.\n\n**Compatibility.** Officially supported: **Claude Code**. Also tested: **Codex** (CLI and the\nChatGPT desktop app — [setup](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#56-using-it-with-codex-tested)) and **opencode**\n([setup](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#57-using-it-with-opencode-tested)). IHMT is a standard stdio MCP server, so any\nagent that supports local MCP servers should work — GitHub Copilot, Antigravity, Cursor, Windsurf,\nGemini CLI, Claude Desktop… — and [`INSTALL.md`](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/INSTALL.md) knows how to configure them, but we have\nnot tested those yet. All of them can share one memory.\n\n| You need | Why | \n|---|---|\n| **Python 3.10 or newer** | IHMT is written in Python | \n| **git** | to download IHMT and keep it updated | \n| **An AI agent that can run commands** (Claude Code, Codex, opencode, Copilot in agent mode…) | it installs IHMT and then uses the memory | \n| **Internet, during the installation** | to download the code and the MCP package; not needed afterwards | \n\n**Missing Python or git? Your AI agent installs them for you** (it is instructed to ask you first).\nPython goes in your user folder, with no administrator password, so nothing system-wide changes. On a brand-new Mac, git may need one click:\nApple shows a window asking to install its command-line tools.\n\nOn a Mac, note that the `python3` that comes with macOS is version 3.9, which is too old — that is why\nyour agent may say Python is missing even though `python3` exists.\n\nYou do **not** need administrator rights, a database, an account, or any paid service beyond your\nagent. Tested on **macOS** and **Linux** (Ubuntu); on **Windows** the instructions are included but\nnot tested yet. Details: [GUIDE.md §3](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#3-requirements).\n\nPaste this into the AI agent you want to give a memory to — **Claude Code**, **Codex** and\n**opencode** are tested; **GitHub Copilot**, **Antigravity**, **Cursor**, **Windsurf**, **Gemini CLI**,\n**Claude Desktop** and other MCP clients should work too:\n\n```\nInstall the IHMT memory MCP server for me from https://github.com/gonzaroman/IHMT-MEMORY — follow the instructions in its INSTALL.md.\n```\n\nThe agent follows [`INSTALL.md`](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/INSTALL.md): it downloads IHMT to `~/IHMT-MEMORY`, keeps your\nmemory in `~/.ihmt`, registers the server **with itself only**, adds the usage instructions, and\ntells you what it did. Then start a new session so the memory tools load.\n\n**Using several agents?** Paste the same prompt in each one, whenever you want. If IHMT is already\ninstalled — say you have used it with Claude for months and now want it in Codex — the agent finds\nthat installation, updates it if it safely can, and connects to the **same memory**, so it knows\nwhat you told the others from day one.\n\nIt needs an agent that can run terminal commands or edit files; chat-only assistants in a browser cannot install anything.\n\n## **Manual install**\n\nRequirements: Python 3.10+, git, and your agent's CLI.\n\n```\ngit clone https://github.com/gonzaroman/IHMT-MEMORY.git ~/IHMT-MEMORY\ncd ~/IHMT-MEMORY\npython3 -m venv .venv\n.venv/bin/pip install -r requirements-mcp.txt\nmkdir -p ~/.ihmt\n```\n\n**Claude Code**\n\n```\nclaude mcp add ihmt-memory -s user -e IHMT_HOME=\"$HOME/.ihmt\" -- \"$PWD/.venv/bin/python\" \"$PWD/mcp_server.py\"\nclaude mcp list                     # ihmt-memory … ✔ Connected\n```\n\nThe server name must come before `-e`. Then append\n[`templates/memory-instructions.md`](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/templates/memory-instructions.md) to `~/.claude/CLAUDE.md`.\n\n**Codex** — `codex mcp add ihmt-memory --env IHMT_HOME=\"$HOME/.ihmt\" -- \"$PWD/.venv/bin/python\" \"$PWD/mcp_server.py\"`,\nthen add `default_tools_approval_mode = \"approve\"` to the `[mcp_servers.ihmt-memory]` table in\n`~/.codex/config.toml` (above its `env` table) and append the template to `~/.codex/AGENTS.md`.\n[Details](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#56-using-it-with-codex-tested).\n\n**opencode** — add an `\"ihmt-memory\"` entry (`\"type\": \"local\"`, `\"command\": [<python>, <mcp_server.py>]`,\n`\"environment\": {\"IHMT_HOME\": <memory folder>}`) to the `\"mcp\"` object of\n`~/.config/opencode/opencode.json`, and append the template to `~/.config/opencode/AGENTS.md`.\n[Details](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#57-using-it-with-opencode-tested).\n\nWindows, the project scope, the graphical setup and troubleshooting are all in the\n[guide](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#4-installation-step-by-step).\n\n**New here? Read [`GUIDE.md`](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md)** — everyday use, step by step, with real outputs. The rest\nof this README is the technical reference.\n\nA universal, domain-agnostic long-term memory for LLMs, stored as a recursive tree of plain files on\nthe local disk. No vector database, no server, **no third-party dependencies** — Python 3.10+ and the\nstandard library.\n\nInstead of embedding everything into one flat index and scanning it, IHMT organizes knowledge into a\ntree: raw text leaves at the bottom, recursive JSON summaries above them, and a single `root.json`\ntrunk at the top. A query walks that tree — root → branch → branch → leaf — so the number of files\nopened grows with the *depth* of the tree (`≈ beam × log_B(n)`), not with the amount stored.\n\n|  | Flat RAG | IHMT | \n|---|---|---|\n| Retrieval cost | scan / ANN over all *n* chunks | `beam × log_B(n)` file reads | \n| Structure | none — a bag of vectors | explicit hierarchy, inspectable | \n| Chunking | fixed character windows | syntax-, scene- and date-aware | \n| Stale facts | served silently | superseded, dated, and flagged | \n| Ambiguity | returns a plausible guess | asks you for a clue | \n| Storage | binary index | UTF-8 `.txt` + JSON you can read | \n\n```\npython3 gui.py                       # graphical interface: set up, browse, inspect\npython3 init_ihmt.py                 # or from the terminal: create ./ihmt_memory\npython3 main.py demo                 # full walkthrough in ./demo_workspace\npython3 -m unittest discover -v      # stdlib only; MCP tests skip without the SDK\n```\n\nThen use it on your own material:\n\n```\npython3 main.py ingest ~/notes ~/project/src/Main.java\npython3 main.py consolidate --force\npython3 main.py search \"how did we handle stock reservations\"\npython3 main.py ask \"Luis\"                    # interactive clue loop\npython3 main.py conflicts                     # what changed over time\n```\n\nAs a library:\n\n``` python\nfrom ihmt import IHMT\n\nmemory = IHMT.initialize(\"./workspace\")\nmemory.ingest_file(\"examples/InventoryService.java\")\nmemory.ingest_file(\"examples/journal_personal.txt\")\nmemory.flush()                                # close the tree up to the root\n\nanswer = memory.search(\"reserveStock soft hold\")\nprint(answer.best.content)                    # the leaf\nprint(answer.best.path)                       # ['root', 'N1-software.java-…', 'L-software.java-…']\nprint(answer.node_reads)                      # how many branch files were opened\n\nfor notice in memory.notices():\n    print(notice)   # \"On 2024-03-11 you said user location = 'Madrid', but on 2026-02-03 you updated to 'Valencia'.\"\npython3 gui.py                       # opens a browser at 127.0.0.1\npython3 gui.py --path ~/my-memory --port 8765 --no-browser\n```\n\nStill zero dependencies — the server is `http.server` from the standard library, it listens only on\nthe loopback interface, and every `/api/*` call needs the random token carried in the URL it opens.\n\nFour screens: **Set up** (pick the memory folder with the system dialog, create the store, choose\nbetween per-project and global registration, preview the exact command or JSON before anything is\nwritten), **Explore** (collapsible tree down to the stored text, with supersession notices),\n**Diagnose** (a search that reports confidence, files opened vs. total, and the descent path), and\n**Timeline** (active vs. historical values and the detected contradictions). The interface is\nbilingual (ES/EN) and read-only over the memory: it never deletes or edits a leaf.\n\n`mcp_server.py` exposes the tree to Claude Code as nine tools in three families: long-term memory,\nproject indexes and session scratch memory. The core stays dependency-free; the\nSDK is an optional extra:\n\n```\npython3 -m venv .venv\n.venv/bin/pip install -r requirements-mcp.txt        # mcp[cli]>=2.0\n```\n\nFor a single project, copy `.mcp.json.example` to that project's `.mcp.json` and fill in the absolute\npaths; Claude Code asks you to approve it on the next session there (`claude mcp list` shows it as\n*Pending approval* until then):\n\n```\n{\n  \"mcpServers\": {\n    \"ihmt-memory\": {\n      \"command\": \"/absolute/path/to/IHMT-MEMORY/.venv/bin/python\",\n      \"args\": [\"/absolute/path/to/IHMT-MEMORY/mcp_server.py\"],\n      \"env\": { \"IHMT_HOME\": \"/absolute/path/to/IHMT-MEMORY\" }\n    }\n  }\n}\n```\n\nor in one command:\n\n```\nclaude mcp add ihmt-memory --scope user \\\n  -e IHMT_HOME=/absolute/path/to/IHMT-MEMORY \\\n  -- /absolute/path/to/IHMT-MEMORY/.venv/bin/python /absolute/path/to/IHMT-MEMORY/mcp_server.py\n```\n\n`IHMT_HOME` selects the store (`$IHMT_HOME/ihmt_memory`), created on first use. Point every project at\none shared directory for a single cross-project memory, or give each project its own.\n\n| Tool | Behaviour | \n|---|---|\n| `search_memory(query, clue=None, detail=\"compact\")` | Walks the tree. Compact output: the best memory with its date and `OUTDATED` notices, one line per other match;`detail=\"full\"` adds ids, tree paths and excerpts. An ambiguous query returns an`AMBIGUOUS` block listing the candidates instead of guessing — call again with`clue` . A query that matches nothing says so, and so does one whose closest entry shares only a stray word with it (`NOT FOUND` ). | \n| `save_memory(content, domain=\"general\", content_type=\"auto\", replaces=\"\")` | Classifies, splits and stores the text, extracts dated facts, and keeps the tree consolidated. Reports how it was filed and — if the save contradicts something remembered earlier — the notice to relay to the user. With `replaces` (a few words describing an earlier memory) the save is recorded as its correction: the old memory is flagged`OUTDATED` and searches answer with the new one first. | \n| `mark_outdated(old_id, new_id)` | Flags one memory as corrected by another, when `save_memory` found several candidates for`replaces` and listed their ids. | \n| `project_map(path, detail=\"files\", subpath=\"\")` | Compact map of a codebase: files with their size in tokens and, with `detail=\"symbols\"` , each method with its line range. Built from the sync manifest, without opening leaves. | \n| `find_code(query, path, scope=\"main\", subpath=\"\", clue=None, max_tokens=1500)` | Returns only the symbol that answers the query, as `file:first-last` + code.`scope` is`main` (skip tests),`test` or`all` . | \n| `read_file(path, force=False)` | Reads a file and remembers what it handed out this session: a repeated read answers `UNCHANGED` or only a unified diff. | \n| `note(text)` /`recall(query, clue=None)` | Session scratch memory: survives a context compaction, disappears when the session ends. | \n| `digest_output(text, label=\"output\")` | Condenses a long log to its first lines, errors, failures, test totals and last lines; the full text stays recallable. | \n\nLong-term memory saves tokens *between* sessions. The project tools save them *within* one, where the\ncost is reading the same files again and again. `ProjectIndex` (`ihmt/project_index.py`) keeps a\nprivate store per project under `$IHMT_PROJECTS_DIR` (default `$IHMT_HOME/ihmt_projects`):\n\n- **one leaf per symbol** —`code_chunk_mode=\"symbol\"` , so a lookup returns a method, not a file;\n- **checksum sync on every call** — size and mtime first, SHA-256 only for what moved; changed files\nare re-ingested, removed ones deleted, and the branches rebuilt. Code is never served stale;\n- **path-ordered branches** — each file is stamped with its rank in path order, so every branch covers\nneighbouring files and its summary stays meaningful for the descent;\n- **code-aware ranking** — the navigator filters by`scope` /`path_prefix` from the catalog, prefers\nthe file a query names, demotes tests unless asked and demotes import lines.\n\nMeasured on a 55-file Spring Boot project (8 typical questions): reading the files that hold the\nanswers costs 3,613 tokens; `find_code` returns the exact method for all 8 in 1,071. The map of the\nproject costs 660 tokens against 13,157 to read it whole.\n\nThose savings are against an agent that **reads whole files**. In an A/B test with 22 real headless\nClaude Code sessions, Claude preferred batched `grep`/` sed -n` and never called the project tools on\nits own; *forcing* them made sessions 42–71 % more expensive. Keeping the server enabled costs about\n460 tokens per conversation, since Claude Code loads MCP tools on demand. IHMT's main value is memory\n**between** sessions; treat the project tools as optional, and do not mandate them in `CLAUDE.md`.\n\nThe server transparently supports MCP SDK 2.x (`MCPServer`), 1.x (` FastMCP`) and the standalone\n`fastmcp` package.\n\nThe usage instructions in your `~/.claude/CLAUDE.md` ([template](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/GUIDE.md#54-tell-claude-when-to-use-it-recommended)) tell Claude Code *when* to reach for each tool: search before answering anything that\ndepends on earlier sessions, save durable facts with their date, never guess on `AMBIGUOUS`, always\nrelay `OUTDATED`, and never store secrets.\n\nEverything lives in one relocatable directory:\n\n```\nihmt_memory/\n  root.json                 # the trunk: domains, topics, top branches, counters\n  layer_0/<domain>/*.txt    # the leaves: raw UTF-8 text + a strict JSON header\n  layers/1/*.json           # branches: summaries of leaves\n  layers/2..N/*.json        # branches: summaries of summaries\n  state/catalog.json        # index: id -> path, domain, parent, timestamp\n  state/facts.json          # the fact timeline\n  ihmt.config.json          # branch factor, token budgets, backend\n```\n\n`root.json` lives *inside* `ihmt_memory/` so the store is self-contained: copy the directory and the\nmemory travels with it.\n\n**A leaf** is a normal text file that describes itself, so it stays meaningful even if the catalog is\nlost:\n\n```\n<<<IHMT-META\n{\n  \"leaf_id\": \"L-software.java-0001-84130ee123\",\n  \"timestamp\": \"2026-09-09T17:45:00Z\",\n  \"data_type\": \"CODE\",\n  \"domain\": \"software.java\",\n  \"tags\": [\"method:reserveStock\", \"class:InventoryService\", \"lang:java\", \"type:code\"],\n  \"parent_id\": \"N1-software.java-5571a7baac\",\n  \"span\": {\"start_line\": 43, \"end_line\": 68},\n  \"checksum\": \"sha256:…\",\n  \"extra\": {\"context\": \"public final class InventoryService {\"}\n}\nIHMT-META>>>\n    public Optional<String> reserveStock(Sku sku, int quantity) {\n        …\n```\n\n**A branch node** embeds each child's title, excerpt and keywords. That is the detail that makes the\ndescent cheap: a branch can be *ranked without opening any of its children*.\n\n`DomainDetector` classifies each document by **type** (`CODE`, `NARRATIVE`, `CLINICAL`, `PERSONAL`,\n`PROCESS`, `GENERIC`) and **domain** (` software.java`, `medicine.clinical`, `process.cooking`,\n`personal`, …) from its extension plus lexical signatures in English and Spanish. Both can be\noverridden with `--domain` / `--type`.\n\nThe type selects the splitter, and every splitter obeys one invariant:\n\n**A logical block is never cut.** If a single block exceeds `max_tokens` it is stored whole and\nflagged `oversized`. Correctness of the block beats hitting the token budget.\n\n- **Code** (`chunkers/code.py` ) — Python via the stdlib`ast` ; Java/JS/TS/C/C++/C#/Go/Rust/Kotlin/\nSwift/PHP via`BraceScanner` , a character-level scanner that tracks brace depth while skipping\ncomments, string literals, char literals, template literals and preprocessor lines. Imports\ncoalesce; each class/function is a block; an oversized class splits**per member** , and the\nenclosing class header travels in the leaf's`extra.context` rather than being spliced into the\ntext. Concatenating a file's leaves reproduces the file**byte for byte** — asserted in the tests.\n- **Narrative** — paragraph-atomic, with`***` ,`---` ,`Chapter` /`Capítulo` as hard boundaries. Only\na paragraph larger than`max_tokens` is split, and then at sentence boundaries.\n- **Temporal** (journals, chats, clinical records) — one dated entry is atomic and a change of date\nis a hard boundary, so a leaf never mixes two encounters.**The date found in the text becomes the\nleaf's timestamp** , which is what makes recency weighting mean*when something was true* rather\nthan when it was ingested.\n- **Process** (recipes, protocols, runbooks) — steps and ingredient lists stay attached to their\nheading.\n\nLeaf ids are derived from `(domain, source, position, content)`, so re-ingesting an unchanged\ndocument rewrites the same leaves instead of duplicating them.\n\nIt watches layer 0. When `branch_factor` leaves of one domain have no parent, it fires a\n**Summarization Event**: they are condensed into a layer-1 node, the node is written, and *only then*\nare the children stamped with their `parent_id` — so an interrupted run re-processes a group instead\nof orphaning it. The same rule applies from layer 1 to layer 2, and so on, until the tree converges;\nthen `root.json` is rewritten.\n\n`consolidate()` is idempotent. `flush()` (`--force`) also promotes partial groups so the tree closes\ncompletely. Anything still unconsolidated is referenced directly by the trunk, so **nothing in the\nstore is ever unreachable from the root**.\n\nRanking uses Okapi BM25 over each candidate's title, keywords, tags and excerpt, with field weights\nand document frequencies computed **across the siblings of the current level** — exactly the\ndiscrimination the descent needs, at no extra I/O cost. The walk keeps a beam of `beam_width`\nbranches per level.\n\nConfidence blends two independent signals:\n\n```\nconfidence = 0.6 × coverage + 0.4 × margin\n```\n\n*Coverage* asks \"does this leaf actually contain what was asked?\"; *margin* asks \"is it\ndistinguishable from its rivals?\". A common first name scores high on the first and near zero on the\nsecond — which is precisely when the system must not guess:\n\n``` bash\n$ python main.py ask \"Luis\"\n\n\"Luis\" is ambiguous (3 memories match this query equally well, confidence 0.62).\nIt could belong to any of these branches:\n  1. [personal] journal_personal.txt · 2024-07-22 — Vacaciones en Benidorm con Luis, mi primo…\n  2. [personal] journal_personal.txt · 2026-08-30 — Fin de semana en la playa de El Saler con Luis…\n  3. [personal] journal_personal.txt · 2024-11-30 — Cierre de trimestre… Luis Marín revisó el pull request…\nGive me a clue to narrow it down (e.g. a place, a date, a project):\n> vacaciones en Benidorm\n\nquery: \"Luis + vacaciones en Benidorm\" · confidence 0.74 · 5 node reads, 3 leaf reads, depth 2\n  1. [personal] journal_personal.txt · 2024-07-22\n     path  root → N2-personal-ea34db33ae → N1-personal-59bf8bc291 → L-personal-0002-a17b3c8f35\n```\n\nThe clue triggers a **joint cross-reference**: candidates matching *both* term groups are boosted\n(×1.6), candidates matching only one are demoted (×0.7). The loop runs up to `max_clue_rounds`\ntimes, stops early if the user declines, and never silently converts an ambiguous query into a\nconfident answer.\n\n`clue_provider` is any callable, so the loop works for a human at a terminal (`input`) or for an\nagent that lets the LLM supply its own follow-up.\n\nFacts are `(subject, attribute, value, timestamp, source_leaf)`, recorded programmatically via\n`record_fact()` or extracted at ingest time by pattern rules (`vivo en X` / `I live in X`,\n`mi stack es Y`, `trabajo en Z`, `Diagnóstico:`, `Tratamiento:`, `Medicación:` …; extend with\n`add_pattern`).\n\nEach `(subject, attribute)` keeps a dated timeline. The newest value is `ACTIVE`; every earlier one\nbecomes `HISTORICAL` with `superseded_by` and a `valid_from`/` valid_to` interval. **Nothing is\ndeleted**, so both questions stay answerable:\n\n```\nmemory.resolver.active_state()[\"user::location\"].value      # 'Valencia'  (now)\nmemory.resolver.state_at(\"2024-12-31\")[\"user::location\"].value  # 'Madrid'  (back then)\n```\n\nRepeating a value at a later date is a confirmation, not a contradiction. A genuine change produces a\ntransparent notice — *\"On 2024-03-11 you said user location = 'Madrid', but on 2026-02-03 you updated to\n'Valencia'.\"* — and the superseded leaf is annotated, so retrieving outdated material always arrives\nwith its correction attached (`SearchResult.notices`).\n\nA leaf itself stays `ACTIVE`: what it says was true *on its own date*, and that remains the right\nanswer to a historical question. What changes is that it can no longer be read as current.\n\n``` python\nclass SummarizerBackend(Protocol):\n    name: str\n    def summarize(self, children, *, domain: str, layer: int) -> NodeSummary: ...\n```\n\n- `HeuristicSummarizer` (default) — stdlib extractive summarization: TF term ranking with EN/ES stop\nwords plus representative-sentence selection. Offline, deterministic, which is what lets the test\nsuite assert on tree shape.\n- `AnthropicSummarizer` (optional) — used only when selected*and* the`anthropic` package and`ANTHROPIC_API_KEY` are both present. Every failure path (missing SDK, missing key, network error,\nunparseable reply) falls back to the heuristic backend, so a consolidation is never lost because a\nmodel was unreachable.\n\n```\npython3 init_ihmt.py --backend anthropic     # model set by summarizer_model in ihmt.config.json\n```\n\nAny other model or local runtime plugs in by implementing the same protocol and passing it as\n`IHMT(..., backend=MyBackend())`.\n\n| Command | Purpose | \n|---|---|\n| `init [--branch-factor N] [--target-tokens N] [--force]` | create the store | \n| `ingest <paths…\\|-> [--domain D] [--type T] [--tag X] [--no-consolidate]` | ingest files, directories or stdin | \n| `consolidate [--force]` | run pending Summarization Events | \n| `search <query> [--top-k N] [--full]` | walk the tree | \n| `ask <query> [--clue TEXT] [--top-k N]` | search with the clue loop | \n| `tree [--depth N]` | outline of the hierarchy | \n| `stats` ,`facts [--subject S]` ,`conflicts [--subject S]` | inspection | \n| `rebuild` | rebuild catalog, timeline and trunk from the files | \n| `demo` | end-to-end walkthrough | \n\n`--path` selects the store directory and `--json` emits machine-readable output; both work before or\nafter the subcommand.\n\n`ihmt_memory/ihmt.config.json`:\n\n| Key | Default | Meaning | \n|---|---|---|\n| `branch_factor` | 8 | children per branch; the log base of retrieval cost | \n| `target_tokens` /`max_tokens` | 2000 / 3000 | leaf size target and oversize threshold | \n| `beam_width` | 3 | branches kept alive per level | \n| `confidence_threshold` | 0.45 | below this, ask for a clue | \n| `ambiguity_margin` | 0.18 | score gap under which candidates count as tied | \n| `max_clue_rounds` | 3 | clue-loop iterations | \n| `summarizer_backend` /`summarizer_model` | `heuristic` /`claude-sonnet-5` | summarization | \n| `code_chunk_mode` | `pack` | `symbol` stores one leaf per class member (used by project indexes) | \n\nSmall corpora deserve a small branch factor — the demo uses `branch_factor=4, target_tokens=400` so a\nhandful of documents still builds a genuine multi-layer tree.\n\n```\npython3 -m unittest discover -v          # from the project root\n```\n\n187 tests — 25 of them for the MCP server, skipped without the SDK — covering: byte-exact reconstruction and boundary-depth invariants for Java and Python, scene/date/section atomicity, oversized-block handling, leaf header round-trips, catalog recovery, summarization thresholds, upward propagation, idempotence, full reachability from the root, descent cost bounds, the clue loop, recency weighting, historical preservation, the project indexes, the MCP tools, the graphical interface and the CLI.\n\n- **Retrieval is a descent, not a scan.** That is the whole point, and it means a branch pruned at\nthe trunk is not revisited. Domain-level pruning only happens when a query has actual signal at the\ntrunk; if it has none, every domain stays in play and the beam applies one level down. The clue\nloop is the recovery mechanism when the descent goes wide.\n- **Lexical, not semantic.** Matching is BM25 over accent-folded, CamelCase-split tokens: it works in\nany language and needs no model, but it will not match a synonym. Plugging an embedding re-ranker\ninto`BM25Ranker` is the natural upgrade; the tree structure does not change.\n- **Token counts are estimated** at ~4 characters per token. Budgets only need to be consistent, not\nexact.\n- **Fact extraction is pattern-based.** The bundled rules cover common English/Spanish phrasings and\nclinical headers;`record_fact()` is the reliable path, and`add_pattern()` extends the rules.\n- **Single-writer.** Writes are atomic (`tmp` +`os.replace` ) and catalog rebuilds take a lock file,\nbut the store assumes one writer at a time.\n- `initialize(force=True)` discards*derived* state only — branches, catalog, timeline — and detaches\nthe surviving leaves so the next consolidation rebuilds the hierarchy. Leaf content is never\ndeleted.\n\n```\nihmt/\n  api.py                    IHMT facade wiring everything together\n  config.py                 IHMTConfig\n  models.py                 MemoryLeaf, BranchNode, ChildRef, RootIndex, Fact, Contradiction\n  storage.py                MemoryStore: atomic I/O, catalog, recovery\n  textutils.py              tokenizing, keywords, entities, extractive summary, timestamps\n  detectors.py              DomainDetector\n  chunkers/                 base · code · narrative · temporal · process · generic\n  summarizers.py            SummarizerBackend · Heuristic · Anthropic\n  universal_ingestor.py     UniversalIngestor\n  recursive_summarizer.py   RecursiveSummarizer\n  semantic_navigator.py     SemanticNavigator, BM25Ranker, ClueRequest, scope/path filters\n  project_index.py          ProjectIndex: per-project code cache with checksum sync\n  conflict_resolver.py      ConflictResolver + timeline manager\nihmt_gui/                   local graphical interface (stdlib only)\nmcp_server.py               MCP server: the nine tools\ninit_ihmt.py · main.py · gui.py · examples/ · tests/\nGUIDE.md                    installation and usage guide\nINSTALL.md                  installation instructions for AI agents\ntemplates/                  memory-instructions.md: the usage rules agents append to CLAUDE.md / AGENTS.md\n```\n\n[MIT](https://github.com/gonzaroman/IHMT-MEMORY/blob/main/LICENSE) © 2026 Gonzalo Román Márquez (gonzaroman)", "url": "https://wpnews.pro/news/long-term-memory-for-ai-coding-agents-as-a-tree-of-plain-files", "canonical_source": "https://github.com/gonzaroman/IHMT-MEMORY", "published_at": "2026-10-11 01:02:14+00:00", "updated_at": "2026-10-11 01:20:43.781884+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "agent-protocols", "ai-tools", "artificial-intelligence"], "entities": ["IHMT-MEMORY", "Claude Code", "Codex", "opencode", "Anthropic", "OpenAI", "Google", "Meta"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/long-term-memory-for-ai-coding-agents-as-a-tree-of-plain-files", "markdown": "https://wpnews.pro/news/long-term-memory-for-ai-coding-agents-as-a-tree-of-plain-files.md", "text": "https://wpnews.pro/news/long-term-memory-for-ai-coding-agents-as-a-tree-of-plain-files.txt", "jsonld": "https://wpnews.pro/news/long-term-memory-for-ai-coding-agents-as-a-tree-of-plain-files.jsonld"}}