This one covers six days, September 30 to October 5. Most of it went into one question for my workflow kit: who runs the tests, and what a green run is allowed to prove. The rest went into an OMEGA release for first-time users, two new brands started on it over the weekend, and the first player tools for the game database.
workkit is my workflow kit for Claude Code: GitHub issues as the work list, labels as the pipeline, and hooks that enforce the rules. It shipped fifteen versions in the window, 0.63.1 to 0.72.0, most of them circling one problem.
At the start of the week the kit ran tests itself, at several gates. Moving an issue to QA reran the touched test files, and the commit gate wanted a full suite run. Then 0.70.0 hashed the tree to skip a rerun when nothing had changed, 0.71.1 read the TAP output instead of guessing from a file's text, and 0.71.2 learned that a test file under a fixtures folder may fail on purpose. Each fix made sense, and the test flow still changed from one day to the next.
0.72.0 turns it around: the kit never runs a test. A package's npm test is the one command, bare for the whole suite or with file arguments for a narrowed run. workkit is set as npm's script shell, so every npm test passes through it. It hashes the working tree, runs the command, and on green writes a record for that hash. That record is the receipt, and the gates only read receipts:
npm test at the root writes the suite record the commit gate checks. workkit prove does the same on a throwaway copy of exactly what is staged.
An agent that can run tests can also run them five times. Tying the proof to a tree hash makes "it passed" a fact about specific bytes, and any edit after the run voids the receipt on its own.
One more change rides on it. A build with a test surface is now a pair of agents briefed from the issue's Contract: one writes the tests, one writes the code, and a verifier proves the new tests fail without the code. That red proof runs in a snapshot taken when the issue was claimed, gitignored build files included, because a copy built from HEAD could crash on a missing file and pass for red.
OMEGA is my framework for shipping a website, backend, desktop app, and browser extension from one project. 0.55.0 shipped on October 1, and its theme is the first ten minutes.
A new user starts from the brand template, a repo with two files. npm start installs OMEGA, runs the onboarding wizard, and boots the site. Both files carry a template marker, so the wizard replaces them on the first run and never on a rerun.
The detail I like: one function, missingEnvKeys, says which env keys a brand still owes, by target, config, and verb. Every verb's env check reads it, so a website-only brand is never asked for a backend key. omega status applies the same idea to a folder: what it is, what it lacks, and the next command. It writes nothing.
Underneath, devkit gained the one test runner every OMEGA package will move to, built on node:test in one process. Desktop and extension switched first, and their suites (1,119 and 351 cases) passed unchanged.
On Saturday two new brands started from that template, the best test onboarding could get.
OddMatter is a creative agency site (private, and the new site is not live yet, so no link). The work was porting a hand-built MVP into an OMEGA web target as its own theme: six pages, 17 sections, and animated effects on one frame loop.
The home hero cost about 40% of a CPU core at rest and stuttered under a CPU slowdown. The frame loop now feeds raw frame times to a small governor with no browser dependency, so a plain Node unit test covers it. When a whole window of frames averages slow, it steps the effects down a level: first the title and the stone slow, then the sky slows and caps its pixel density. It never steps back up during a page view, so the page does not flicker between looks.
Suponsa is a marketplace for guest post and link insertion slots (private, no live site yet). Its first real commit is the submission flow: buyers submit through a wizard, and sellers approve, decline, ask for changes, mark paid, and deliver. The status machine lives in the backend, and a submission stores its own next moves, so a page renders without reading the seller's documents. Each move emails the other side after the write, and a failed email never fails the move.
Both scaffolds also turned up framework gaps, which went back to OMEGA as issues.
Lootlore, the game database from the last post (still private, no public URL), got three tool pages: an EXP table, a training spot finder, and a scroll simulator.
The training page ranks hunting maps by EXP per hour. Attack time comes from the frames of the basic attack at the worn weapon's speed tier, and kills per hour from each map's spawn supply for one player, using formulas a community site measured. The first version ranked by EXP per swing, which ignores both how fast you swing and how crowded a map is.
Skill pages also open with a small live scene: a character of that class casts the skill in the game's own art, with the game's damage numbers.
singleInstance switch OMEGA added for CLI-shaped desktop apps. A notification is never handed to a running copy that may have hung.
Next: the rest of OMEGA's packages move to the one test runner, and the two new brands head for their first deploy. If your agents or your CI decide when tests run, I would like to hear how you keep them from running twice. Comments open.