A Harness in a single file: Agent skills as an evolving REPL module A developer built DeepClause, a Prolog-dialect agent harness that runs inside a REPL, letting an agent dynamically add, execute, and modify DML predicates until it solves a task. The approach aims at recursively self-improving (RSI) harnesses that can be versioned, exported, and shared as a single file, with transactional rollback of code changes in the REPL state. tldr; An Agent plus a REPL let’s us build RSI/RLM style harnesses that can evolve over time, get versioned and easily be exported, shared and imported - if the harness is defined as a set of predicates in a Prolog-like language or possibly Lisp ;- . As usual, this is written by a human, code/AI-slop is on GitHub https://github.com/deepclause/deepclause-sdk ;- For a while now I’ve been somewhat bothered by the fact that DeepClause’s DML Language https://github.com/deepclause/deepclause-sdk doesn’t come with a REPL. While I was thinking a bit about how to add one, I quickly realized that it would be much more interesting to build a combined Agent/REPL kind of thing - especially in the context of RSI, meaning “recursively self improving” Harnesses See references at the end . So the idea became: 1. Build an agent that has only access to a REPL that can load and execute DML code Plus some tools to load reference docs into context and inspect the current REPL state . 2. Instruct the agent to use its access to the REPL to dynamically add, execute, and modify code until it can successively solve a task. DML Deepclause Meta Language is a Prolog dialect specifically designed to express mixed deterministic and LLM-based workflows and agents. Given its feature set, there are a few points here that make DML and its Prolog ancestry very useful in the context of RSI harnesses: 1. State management: Prolog is “functional” actually it’s more “relational”, but that’s another story… and generally stateless. The loaded predicates ”functions” define the entire state and can be easily inspected and modified on the fly. There are no global variables or references in the interpreter memory. 2. RLM-like structures: Code written by the agent will be much more useful if we can adopt an RLM-like structure, that means that the code can again call or orchestrate underlying LLM models or sub-agents. This is exactly what DML was designed for in the first place We have predicates that trigger entire agentic loops, call decision models or run a single prompt with fresh context. At the same time we have access to host tools such as e.g. bash and can reach them through DML code. 3. Transactional changes: Another very interesting point is that the simplicity of the REPL state management allows one to easily implement transactions around changes to the code in the REPL. That gives the code writing agent an easy way to test and potentially rollback changes without polluting the code and REPL state too much. An RLM-style harness defined in a single file So how does it work? First, install deepclause via npm, configure model/provider and start the REPL: npm install -g deepclause-sdk deepclause init deepclause login deepclause repl You will then be greeted with a banner and a running REPL session: Now you can directly enter prompt or execute a DML/Prolog goal using e.g. “/goal member X, a,b,c,d .” Much more interesting is of course to start with a prompt. So let’s enter “Build me a small coding agent ” and you will see output like this: Eventually, if you use a decent model, you can inspect what the resulting harness looks like using “/listing” and you will see code like this module doc "Coding agent: read, search, and edit a repo and run its tests." . % --- data = code: facts the agent can query and reuse --- doc project/1, "Known projects in this workspace." . project payments . project checkout . % --- tools: predicates, plus a tool/2 annotation that exposes them to task/N --- run shell Command, Output :- exec bash command: Command , Result , get dict stdout, Result, Out - Output = Out ; Output = "" . fs read Path, Text :- read file Path, Text . fs write Path, Content :- write file Path, Content . tool run shell Command, Output , "Run a shell command; returns stdout." :- run shell Command, Output . tool fs read Path, Text , "Read a workspace file." :- fs read Path, Text . tool fs write Path, Content , "Write a workspace file." :- fs write Path, Content . % --- skills: predicates. This one scopes the tools a model turn may call. --- doc patch/3, "Apply a minimal fix and run the tests. patch +File, +Issue, -Summary ." . patch File, Issue, Summary :- system "Make the smallest change that fixes the issue. Run the project's tests." , with tools run shell, fs read, fs write , task "Fix this issue in {File}: {Issue}. Then run the tests and report what changed.", Summary . % --- memory: the module can also read its own past --- doc recent changes/2, "The last N committed changes." . recent changes N, Changes :- findall V-Ts, harness change V, Ts, commit, , , All , last n All, N, Changes . This is a snapshot of all the functionality that the harness provides at this point. The various parts of this fall into one of these categories: 1. Descriptions/Annotations/Instructions, e.g. all the things you would normally spread across a system prompt, a tool schema, and a separate SKILL.md live in the module. Note the module doc and doc predicates which merely act as descriptions of how this harness should be used. 2. Ordinary Code: File I/O, network requests to services for, arithmetic, so anything that would be handled by tools. A harness may also need some external scripts e.g. a Python script to query an API . They can be linked using script file predicates, so that the system is aware of them. 3. Sub-Agents: If you look at the “patch” predicate you can see that it models an agent: system prompt, an agentic loop task using a set of tools with tools . And that’s exactly what’s we have here. When the REPL/Agent decides to run this subagent or rather just execute the predicate it first starts a new execution frame with fresh context by default , builds context as defined in the above code and starts the agent loop. 4. Memory: Prolog facts can be used to store almost any data along side the harness and make it directly queryable through the REPL or a natural language prompt Tables, mappings, histories, “Who is Alice’s uncle’s second cousin?”, etc. . So now that we have all that in a single file, it’s very easy to share it or commit and version it somewhere else. Exporting and importing a harness Using “/export” we can serialize the current module state a plain, self-contained dml file. The export is deterministic: directives first, then tools, then embedded scripts, then user predicates sorted by name/arity. If a file links to scripts python/bash using a script file predicate, then the whole harness is exported as a zip file with those linked scripts. To import and re-use it: deepclause repl -l coding agent.dml load and continue interactively deepclause repl -l coding agent.dml -f load frozen read-only deepclause run coding agent.dml --goal ‘patch ”src/cli.ts”, “…”, S ’ Advantages and use cases Why bother with a single-file harness instead of a prompt plus a pile of skills? - Model-specific harnesses: Different models need different amounts of scaffolding. A strong model can be given a fairly autonomous loop; a small or local model does much better with a short, explicit workflow and a single model call at the end. Since the harness is text, you can keep a tuned harness per model and compare them on the same tasks. - Task-specific harnesses: Only load the harness that is tailored to a task, other skills and extensions for your coding agent won’t confuse your model. - Continuous optimization and RSI: In the above framework, a harness becomes a mutable artifact with a version history and a rollback. You can change one predicate, test it, measure it against a small task set, and keep the change only if it helps. This is exactly the edit → verify → keep loop the recent self-improvement literature appears to converge on see references , with the difference that the object being edited is one readable file. - Cost, latency and determinism: Not everything needs to be handled by a full agent or an LLM, but can be implemented using plain old if-then-else or for/while. Also, for some stuff we’d might want to use decision models like Jev which DeepCLasue also supports or other more traditional approaches. - Auditability and review. Every change is a diff with a version and every capability is a implemented in one or more predicates. The tools a model may use can easily be restricted using “with tools/2”. Execution of any DML code gets tracked by a mtea-interpreter, so a full trace can be generated. This makes it very useful for regulated industry use cases. - Portability and sharing. One file, no registry, import, change, re-export. Where to take it from here? Well the answer is obvious: enter the prompt “Please iterate on yourself until you become sentient.” Jokes aside, if you’re interested or would like to discuss further, please do feel free to try out the latest version of DeepClause 0.1.0 and/or leave a comment with any suggestions for what one could build with it References this bit was compiled by DeepSeek : - MESH-Harness https://arxiv.org/abs/2610.05300 treats the harness as “the code that organizes context, maintains state, and coordinates tool calls,” and improves it by swapping and recombining explicit modules under a fixed evaluation budget. - The Harness as the Only Mutable Surface https://arxiv.org/abs/2610.10629 argues that self-evolution is reviewable only when it is confined to the harness, so that “every adaptation is a diff with a cause and a test attached.” - SAGE https://arxiv.org/abs/2609.36043 frames the loop as an optimizer that proposes an edit to a persistent skill document plus a gate that accepts or rejects it — and shows why a naive gate is not enough. - AgentEvolver https://arxiv.org/abs/2610.11613 , Skill-V https://arxiv.org/abs/2610.11781 and SkillForge https://arxiv.org/abs/2610.09832 all turn experience into reusable capability with a versioned, auditable lifecycle. - SelfSearch https://arxiv.org/abs/2609.37968 has agents inspect and modify their own instructions, tools and execution procedures.