{"slug": "a-harness-in-a-single-file-agent-skills-as-an-evolving-repl-module", "title": "A Harness in a single file: Agent skills as an evolving REPL module", "summary": "A developer built DeepClause, a Prolog-dialect agent harness that runs inside a REPL, letting an agent dynamically add, execute, and modify DML predicates until it solves a task. The approach aims at recursively self-improving (RSI) harnesses that can be versioned, exported, and shared as a single file, with transactional rollback of code changes in the REPL state.", "body_md": "**tldr; An Agent plus a REPL let’s us build RSI/RLM style harnesses that can evolve over time, get versioned and easily be exported, shared and imported - if the harness is defined as a set of predicates in a Prolog-like language (or possibly Lisp ;-).** \n\n*(As usual, this is written by a human, code/AI-slop is on [GitHub](https://github.com/deepclause/deepclause-sdk) ;-)*\n\nFor a while now I’ve been somewhat bothered by the fact that [DeepClause’s DML Language](https://github.com/deepclause/deepclause-sdk) doesn’t come with a REPL. While I was thinking a bit about how to add one, I quickly realized that it would be much more interesting to build a combined Agent/REPL kind of thing - especially in the context of RSI, meaning **“recursively self improving” Harnesses** (See references at the end).\n\nSo the idea became:\n\n1. Build an agent that has only access to a REPL that can load and execute DML code (Plus some tools to load reference docs into context and inspect the current REPL state).\n2. Instruct the agent to use its access to the REPL to dynamically add, execute, and modify code until it can successively solve a task.\n\nDML (Deepclause Meta Language) is a Prolog dialect specifically designed to express mixed deterministic and LLM-based workflows and agents. Given its feature set, there are a few points here that make DML (and its Prolog ancestry) very useful in the context of RSI harnesses:\n\n1. **State management:** Prolog is “functional” (actually it’s more “relational”, but that’s another story…) and generally stateless. The loaded predicates (”functions”) define the entire state and can be easily inspected and modified on the fly.  There are no global variables or references in the interpreter memory.\n2. **RLM-like structures:** Code written by the agent will be much more useful if we can adopt an RLM-like structure, that means that the code can again call or orchestrate underlying LLM models or sub-agents. This is exactly what DML was designed for in the first place! We have predicates that trigger entire agentic loops, call decision models or run a single prompt with fresh context. At the same time we have access to host tools (such as e.g. bash) and can reach them through DML code.\n3. **Transactional changes:** Another very interesting point is that the simplicity of the REPL state management  allows one to easily implement transactions around changes to the code in the REPL. That gives the code writing agent an easy way to test and potentially rollback changes without polluting the code and REPL state too much.\n\n**An RLM-style harness defined in a single file** \n\nSo how does it work? First, install deepclause via npm, configure model/provider and start the REPL:\n\n```\n#> npm install -g deepclause-sdk\n#> deepclause init\n#> deepclause login\n#> deepclause repl\n```\n\nYou will then be greeted with a banner and a running REPL session:\n\nNow you can directly enter prompt or execute a DML/Prolog goal using e.g. “/goal member(X, [a,b,c,d]).”\n\nMuch more interesting is of course to  start with a prompt. So let’s enter **“Build me a small coding agent!”** and you will see output like this:\n\nEventually, if you use a decent model, you can inspect what the resulting harness looks like using “/listing” and you will see code like this\n\n```\nmodule_doc(\"Coding agent: read, search, and edit a repo and run its tests.\").\n\n% --- data = code: facts the agent can query and reuse ---\ndoc(project/1, \"Known projects in this workspace.\").\nproject(payments).\nproject(checkout).\n\n% --- tools: predicates, plus a tool/2 annotation that exposes them to task/N ---\nrun_shell(Command, Output) :-\n    exec(bash(command: Command), Result),\n    ( get_dict(stdout, Result, Out) -> Output = Out ; Output = \"\" ).\n\nfs_read(Path, Text) :- read_file(Path, Text).\nfs_write(Path, Content) :- write_file(Path, Content).\n\ntool(run_shell(Command, Output), \"Run a shell command; returns stdout.\") :- run_shell(Command, Output).\ntool(fs_read(Path, Text), \"Read a workspace file.\") :- fs_read(Path, Text).\ntool(fs_write(Path, Content), \"Write a workspace file.\") :- fs_write(Path, Content).\n\n% --- skills: predicates. This one scopes the tools a model turn may call. ---\ndoc(patch/3, \"Apply a minimal fix and run the tests. patch(+File, +Issue, -Summary).\").\npatch(File, Issue, Summary) :-\n    system(\"Make the smallest change that fixes the issue. Run the project's tests.\"),\n    with_tools([run_shell, fs_read, fs_write],\n        task(\"Fix this issue in {File}: {Issue}. Then run the tests and report what changed.\", Summary)).\n\n% --- memory: the module can also read its own past ---\ndoc(recent_changes/2, \"The last N committed changes.\").\nrecent_changes(N, Changes) :-\n    findall(V-Ts, harness_change(V, Ts, commit, _, _), All),\n    last_n(All, N, Changes).\n```\n\nThis is a snapshot of all the functionality that the harness provides at this point. The various parts of this fall into one of these categories:\n\n1. **Descriptions/Annotations/Instructions,** e.g. all the things you would normally spread across a system prompt, a tool schema, and a separate SKILL.md live in the module. Note the module_doc and doc predicates which merely act as descriptions of how this harness should be used.\n2. **Ordinary Code:** File I/O, network requests to services for, arithmetic, so anything that would be handled by tools. A harness may also need some external scripts (e.g. a Python script to query an API). They can be linked using script_file predicates, so that the system is aware of them.\n3. **Sub-Agents:** If you look at the “patch” predicate you can see that it models an agent: system prompt, an agentic loop (task) using a set of tools (with_tools). And that’s exactly what’s we have here. When the REPL/Agent decides to run this subagent (or rather just execute the predicate) it first starts a new execution frame (with fresh context by default), builds context as defined in the above code and starts the agent loop.\n4. **Memory:** Prolog facts can be used to store almost any data along side the harness and make it directly queryable through the REPL or a natural language prompt (Tables, mappings, histories, “Who is Alice’s uncle’s second cousin?”, etc.).\n\nSo now that we have all that in a single file, it’s very easy to share it or commit and version it somewhere else.\n\n**Exporting and importing a harness**\n\nUsing “/export” we can serialize the current module state a plain, self-contained dml file. The export is deterministic: directives first, then tools, then embedded scripts, then user predicates sorted by name/arity. If a file links to scripts (python/bash) using a script_file predicate, then the whole harness is exported as a zip file with those linked scripts.\n\nTo import and re-use it:\n\n```\n#> deepclause repl -l coding_agent.dml          # load and continue interactively\n#> deepclause repl -l coding_agent.dml -f       # load frozen (read-only)\n#> deepclause run coding_agent.dml --goal ‘patch(”src/cli.ts”, “…”, S)’\n```\n\n**Advantages and use cases**\n\nWhy bother with a single-file harness instead of a prompt plus a pile of skills?\n\n- **Model-specific harnesses:** Different models need different amounts of scaffolding. A strong model can be given a fairly autonomous loop; a small or local model does much better with a short, explicit workflow and a single model call at the end. Since the harness is text, you can keep a tuned harness per model and compare them on the same tasks.\n- **Task-specific harnesses:** Only load the harness that is tailored to a task, other skills and extensions for your coding agent won’t confuse your model.\n- **Continuous optimization and RSI:** In the above framework, a harness becomes a mutable artifact with a version history and a rollback. You can change one predicate, test it, measure it against a small task set, and keep the change only if it helps. This is exactly the**edit → verify → keep** loop the recent self-improvement literature appears to converge on (see references), with the difference that the object being edited is one readable file.\n- **Cost, latency and determinism:** Not everything needs to be handled by a full agent or an LLM, but can be implemented using plain old if-then-else or for/while. Also, for some stuff we’d might want to use decision models like Jev (which DeepCLasue also supports) or other more traditional approaches.\n- **Auditability and review.** Every change is a diff with a version and every capability is a implemented in one or more predicates. The tools a model may use can easily be restricted using “with_tools/2”.  Execution of any DML code gets tracked by a mtea-interpreter, so a full trace can be generated. This makes it very useful for regulated industry use cases.\n- **Portability and sharing.** One file, no registry, import, change, re-export.\n\n**Where to take it from here?**\n\nWell the answer is obvious: enter the prompt “Please iterate on yourself until you become sentient.”\n\nJokes aside, if you’re interested or would like to discuss further, please do feel free to try out the latest version of DeepClause (0.1.0) and/or leave a comment with any suggestions for what one could build with it!\n\n**References (this bit was compiled by DeepSeek):**\n\n- MESH-Harness (https://arxiv.org/abs/2610.05300) treats the harness as “the code that organizes context, maintains state, and coordinates tool calls,” and improves it by swapping and recombining explicit modules under a fixed evaluation budget.\n\n- The Harness as the Only Mutable Surface** (https://arxiv.org/abs/2610.10629) argues that self-evolution is reviewable *only* when it is confined to the harness, so that “every adaptation is a diff with a cause and a test attached.”\n\n- SAGE (https://arxiv.org/abs/2609.36043) frames the loop as an optimizer that proposes an edit to a persistent skill document plus a *gate* that accepts or rejects it — and shows why a naive gate is not enough.\n\n- AgentEvolver (https://arxiv.org/abs/2610.11613), **Skill-V**(https://arxiv.org/abs/2610.11781)and **SkillForge** (https://arxiv.org/abs/2610.09832) all turn experience into reusable capability with a versioned, auditable lifecycle.\n\n- SelfSearch (https://arxiv.org/abs/2609.37968) has agents inspect and modify their own instructions, tools and execution procedures.", "url": "https://wpnews.pro/news/a-harness-in-a-single-file-agent-skills-as-an-evolving-repl-module", "canonical_source": "https://deepclause.substack.com/p/a-harness-in-a-single-file-agent", "published_at": "2026-10-09 15:45:17+00:00", "updated_at": "2026-10-09 15:54:39.939682+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models"], "entities": ["DeepClause", "DML", "Prolog", "GitHub", "npm"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-harness-in-a-single-file-agent-skills-as-an-evolving-repl-module", "markdown": "https://wpnews.pro/news/a-harness-in-a-single-file-agent-skills-as-an-evolving-repl-module.md", "text": "https://wpnews.pro/news/a-harness-in-a-single-file-agent-skills-as-an-evolving-repl-module.txt", "jsonld": "https://wpnews.pro/news/a-harness-in-a-single-file-agent-skills-as-an-evolving-repl-module.jsonld"}}