{"slug": "show-hn-recast-an-experimental-smalltalk-app-that-repairs-itself-with-an-llm", "title": "Show HN: Recast, an experimental Smalltalk app that repairs itself with an LLM", "summary": "Developer \"rot13maxi\" released Recast, an experimental Smalltalk application that wires an LLM into a running Pharo 12 image's exception path so a failing method call can be repaired and hot-swapped without a rebuild or restart. In the demo, a vending machine's deliberately broken `dispense` method fails on the first call, then a separate Pharo VM verifies the model's proposed patch and the image applies it, so the next call returns 'product'; the project ships a GUIDE.md tour and an offline fake LLM path that needs no API key. Recast's author frames the work as an experiment in how much of an application's maintenance loop can live inside the application itself, noting the demo covers one class, method replacements, simple literal receiver state, and two application tests rather than general autonomous software maintenance.", "body_md": "**A Smalltalk application that uses an LLM to repair itself while it runs.**\n\nRecast wires an LLM into a running application's exception path. A failing call supplies the stack, method source, receiver state, and tests. The model can propose a repair; a separate VM checks it before the application replaces the method in its live image. Subsequent calls use the new code, without a rebuild or restart.\n\nThe working demo is a vending machine with a deliberately broken `dispense`\nmethod. With automatic repair enabled, an ordinary call triggers the whole\nloop—no `lk heal` command or manually written repair prompt. Here is a\nshortened transcript, after injecting the failure:\n\n```\n$ ./lk eval \"| vm | vm := VendingMachine make. vm insertCoin: 60. vm dispense\"\nERROR: Error: automatic repair demo\n\n$ ./lk incidents\n... \"status\": \"awaiting-model\" ...\n... \"status\": \"verifying\" ...\n... \"status\": \"repaired\", \"detail\": { \"ok\": true, \"repro\": \"true\", ... }\n\n$ ./lk eval \"| vm | vm := VendingMachine make. vm insertCoin: 60. vm dispense\"\n'product'\n```\n\nThe first call still fails. Investigation and verification run asynchronously; the image applies the accepted patch, and the next call succeeds. This path has been exercised with a real LLM on Pharo 12.\n\n**[Take the tour in GUIDE.md](https://github.com/rot13maxi/recast/blob/main/GUIDE.md)** to reproduce it: boot, edit live code,\nbreak the app, trigger automatic repair, and promote the result. You can also\nuse the [offline fake LLM](https://github.com/rot13maxi/recast/blob/main/GUIDE.md#offline-path-use-the-fake-llm) without an\nAPI key; it supplies canned proposals through the same verification path.\n\nAn application's runtime has useful evidence about a failure: the call stack, the receiver's state, and the code that actually ran. Recast turns that evidence into a repair request and connects the result to a live deployment mechanism. The experiment is how much of an application's maintenance loop can become part of the application itself.\n\nSmalltalk is a useful substrate because classes, methods, and execution contexts are objects the running system can inspect. Compile a replacement method into a class, and subsequent message sends use it immediately. The existing application objects keep their identity; applying a method patch does not require rebuilding their world.\n\nRecast adds a verification and persistence policy around that capability: test the proposal in another VM, apply it tentatively, and explicitly promote the definitions that should survive restart. The LLM proposes; the runtime checks and applies.\n\nThe example is intentionally small: one demo class, method replacements, receivers with simple literal state, and two application tests. It demonstrates the mechanics of automatic repair, not general autonomous software maintenance. A green test run establishes only what those tests check.\n\n``` php\nflowchart LR\n    Failure[Application call fails] --> Capture[Capture incident]\n    Capture --> LLM[LLM investigates]\n    LLM --> Decision{Repair justified?}\n    Decision -->|Yes| Shadow[Separate VM: repro and tests]\n    Decision -->|No| Audit[Record decision]\n    Shadow -->|Pass| Live[Hot-swap live method]\n    Shadow -->|Fail| Audit\n    Live --> Promote[Explicit promotion]\n```\n\n1. **Capture.**`LLMRepairService` observes an error at the application's eval\nboundary, before it becomes an error string. For an opted-in application\nclass, it captures the failing method, stack, source, simple receiver state,\narguments, existing tests, and the class's repair contract. It constructs a\nrepro using a fresh receiver with the captured state.\n2. **Investigate.** The Python supervisor sends that evidence to the model.\nThe model returns`expected` ,`insufficient-evidence` , or`repair` , with a\nreason. A raised exception alone does not require a patch.\n3. **Verify.** The image parses a repair proposal and permits only a replacement\nof the captured application method. A separate Pharo VM loads the kernel,\nreplays the journal, applies the candidate, and runs the repro and suite.\n4. **Apply.** A passing verdict lets the live image compile the replacement and\njournal it. The failed operation is never automatically retried. Promotion\nremains a separate operator action.\n\nThe application queue and heartbeat continue while the model and shadow VM\nwork. Repeated failures from an unchanged method are deduplicated. A code edit\ninvalidates an outstanding proposal; disabling repair or restarting cancels\npending work. `lk incidents` records decisions and verification results, and\n`data/heal/req-N.json` / `resp-N.json` preserve the captured evidence and replies.\n\nAutomatic repair defaults to off and currently opts in `VendingMachine`.\nExpected insufficient-credit errors, compilation errors, test runs, boot\nreplay, and shadow evaluation do not start investigations. Another application\nclass can opt in through `repairClasses` in kernel configuration and a\n`repairContract` class method; `repairExpectedErrors` lists error messages to\nexclude.\n\nYou need Git, Docker with the Compose plugin, Python 3, and an API key for an OpenAI-compatible Chat Completions endpoint. From a fresh checkout:\n\n```\ngit clone https://github.com/rot13maxi/recast.git\ncd recast\ncp .env.example .env\n# Edit .env: set your endpoint, model, and LIVE_LLM_API_KEY.\ndocker build -f Dockerfile.base -t live-smalltalk-base:latest .\ndocker compose up -d --build\n./lk status\n./lk tests\n```\n\nThe template includes an OpenAI example endpoint and model. You can configure\nanother compatible provider or a local model. The GUIDE explains the request\nformat and includes a connection check from inside the container. Real-model\nrepair runs have been validated; check your chosen endpoint before starting\nthe tour. If `.env` already exists, edit it instead of overwriting it. It is\nignored by Git. The first build downloads Pharo and its dependencies.\n\nContinue with [the GUIDE](https://github.com/rot13maxi/recast/blob/main/GUIDE.md) for the failure-to-repair experiment.\nThe command-line interface also exposes each part of the loop:\n\n| Command | Purpose | \n|---|---|\n| `./lk eval \"...\"` | Evaluate code or compile definitions directly in the live VM | \n| `./lk autorepair on` /`off` | Enable or disable incident-driven repair | \n| `./lk incidents` | Inspect automatic investigations and their outcomes | \n| `./lk candidate fix.st --repro repro.st` | Verify a patch without applying it | \n| `./lk heal --problem \"...\" --repro repro.st --target VendingMachine` | Request a repair explicitly | \n| `./lk tests` | Run the registered application tests | \n| `./lk promote` | Make journalled definitions survive restart if tests pass | \n| `./lk rollback 2` | Set the journal prefix to replay on the next boot | \n\nFor explicit `lk heal`, the model receives the supplied problem and repro\nwithout the automatic incident's source and state capture. Include the\nrelevant fields and intended behavior in the problem description. This path\nis an [optional experiment in the GUIDE](https://github.com/rot13maxi/recast/blob/main/GUIDE.md#5-optional-request-a-repair-explicitly).\n\nThe live image never saves a heap snapshot. Each boot loads the base image,\ninstalls the kernel, replays the promoted journal prefix, and starts the queue.\nMutable files live in `/data`, bind-mounted from the host's `data/` directory:\n\n| File | Purpose | \n|---|---|\n| `journal.json` | Ordered sources for successful definition edits | \n| `marker.json` | `{n}` : the first`n` journal entries are promoted | \n| `config.json` | Kernel settings, including `automaticRepair` and`repairClasses` | \n| `tests.json` | Registered test class names | \n| `incidents.json` | Automatic investigations and their outcomes | \n| `healed.json` | Successful repairs | \n\nDefinitions beyond the marker are tentative. Restart leaves them unapplied;\n`lk promote` runs the suite and advances the marker to the journal size if it\npasses. `lk rollback <n>` chooses an earlier prefix for the next boot.\nThe journal restores definitions, not application data or in-memory objects.\n\nReal-model runs exercised both automatic and explicit repair. The proposed vending-machine fixes passed the shadow repro and application suite and were applied live. Promotion and restart preserved a repair; an unpromoted edit was left unapplied after restart.\n\nDeterministic integration checks cover rejected candidates, patches outside the allowed method, stale proposals, duplicate suppression, expected failures, declined investigations, and cancellation. A deliberately slow shadow run leaves the application queue and heartbeat responsive. The offline endpoint returns predefined patches; those checks test the mechanism, not model quality. There is no repair success-rate benchmark here.\n\nTo rerun the checks against a demo stack:\n\n```\npython3 -m unittest discover -s test -p 'test_*.py'\ndocker compose stop supervisor\npython3 test/auto_repair.py\ndocker compose start supervisor\n```\n\nThe integration script makes tentative edits, restores the vending method and automatic-repair setting, and leaves its journal entries and audit trail.\n\n- **Verification depends on tests.** The demo has two application tests. The\nmodel gets an application contract, but that prose is not an independent\nexecutable verifier. Stronger tests and independent specifications matter.\n- **The shadow VM is not a sandbox.** Candidate code can access shared files\nand runtime facilities. Restricting which method can be replaced does not\nconstrain everything its body can do. Hostile proposals need stronger\nprocess and filesystem isolation.\n- **Repair does not undo side effects.** Receiver state is captured at failure\ntime. The failed call may already have changed state, so it is never blindly\nretried. Complex object graphs are skipped.\n- **Method repair is narrower than state migration.** Changing instance-variable\nlayouts or assumptions about existing objects needs migration machinery.\nPassing tests in a fresh VM does not establish the live heap's compatibility.\n- **Replay is code persistence.** Application data needs separate storage.\n\nThe next experiments are stronger isolation, independent tests, canary traffic, and explicit support for state migration and operations that can be retried.\n\n| File | Role | \n|---|---|\n| `kernel/gen_methods.py` | Generates the kernel sources and installer | \n| `kernel/repair_methods.py` | Incident capture and automatic repair state machine | \n| `kernel/methods.json` | Generated Smalltalk method sources | \n| `kernel/boot_eval.st` | Kernel installation, replay, and queue startup | \n| `agent/supervisor.py` | LLM requests and investigation decisions | \n| `scripts/shadow_run.sh` | Candidate evaluation in a separate VM | \n| `lk` | Host CLI using the file queue | \n| `test/auto_repair.py` | Live integration checks with the supervisor stopped | \n| `test/test_supervisor.py` | Response parsing and publication checks | \n| `test/fakellm.py` | Offline OpenAI-compatible fixture server | \n| `Dockerfile.base` /`Dockerfile` /`docker-compose.yml` | Pharo base image and demo stack | \n\nWhen extending the kernel, edit the Python sources and run\n`python3 kernel/gen_methods.py`. The installer creates classes before compiling\nmethods; method sources are stored as JSON strings to avoid manual quote\nescaping. Boot and shadow replay avoid live-side effects such as incident\ncapture and automatic test registration.", "url": "https://wpnews.pro/news/show-hn-recast-an-experimental-smalltalk-app-that-repairs-itself-with-an-llm", "canonical_source": "https://github.com/rot13maxi/recast", "published_at": "2026-10-05 15:08:03+00:00", "updated_at": "2026-10-05 15:20:56.286387+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "developer-tools"], "entities": ["Recast", "Smalltalk", "Pharo 12", "LLMRepairService", "rot13maxi"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-recast-an-experimental-smalltalk-app-that-repairs-itself-with-an-llm", "markdown": "https://wpnews.pro/news/show-hn-recast-an-experimental-smalltalk-app-that-repairs-itself-with-an-llm.md", "text": "https://wpnews.pro/news/show-hn-recast-an-experimental-smalltalk-app-that-repairs-itself-with-an-llm.txt", "jsonld": "https://wpnews.pro/news/show-hn-recast-an-experimental-smalltalk-app-that-repairs-itself-with-an-llm.jsonld"}}