cd /news/artificial-intelligence/the-pit-crew · home topics artificial-intelligence article
[ARTICLE · art-110976] src=thewatershed.markpesce.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

The Pit Crew

Mark Pesce, a technology writer, reported that he used AI agents to optimize his local Qwen3.8-27B model, achieving a 20% speed improvement after GPT-5.6 Sol in Codex spent four hours tuning it. He also used the anonymous Ox Alpha model to draft configuration changes for the Hermes harness, which were applied by GPT-5.6 Terra, demonstrating a workflow where frontier models act as a 'pit crew' for home AI systems.

read3 min views2 publishedAug 25, 2026
The Pit Crew
Image: Thewatershed (auto-discovered)

I've been having too much fun with the new home watershed AI models. But the cutting edge sometimes draws blood, and on Monday I had to disentangle a model from a task that it simply could not complete. Caught in a loop, it repeated the same mistake, over and over again.

I'd heard of such things happening. I'd never seen it.

I pulled back, raised the quality settings, let it go back to work. No change.

Sometimes the knife you think is 'sharp enough' turns out to be rather dull.

So, back into the drawer for the old dependable: Qwen3.8-27B. (Eleven days old, which seems like a decade this month.) I'd put it aside for two reasons: it's very slow, and because it's very slow it can get "caught" in the Hermes harness I use it within.

Hermes is fantastic, but it has some settings that make it less than useful for long-horizon tasks. By this I mean tasks that run on the nightshift - from when I go to bed to when I wake up again. If a harness can't hold itself together for an eight-hour task, it's only getting in the way of the agent.

The speed issue is a classic optimisation problem, and the kind that's perfect to hand to a very smart AI agent - like GPT-5.6 Sol inside Codex. I told Sol to make my model fast and accurate and reliable, favouring accuracy and reliability over speed. It spent nearly four hours running a series of tests, taking benchmarks, tweaking settings, and testing again. At the end of all of that I have Qwen running 20% faster than it was before. Not insignificant - though I am fighting back envy of folks with GPUs big enough to run the model at 30 or even 130 tps. Luxury!

Then onto Hermes, designed for agentic work of a more interactive variety than found on the unattended nightshift. I loaded up Ox Alpha in Hermes - an anonymous free model that's stunned users with its power, and raised some big questions about its origins - and asked it to draft a set of proposed changes to Hermes' configuration to accommodate my slow model chugging its way through the nightshift. I handed off those recommendations to Codex and GPT-5.6 Terra, because Hermes agents aren't allowed to manipulate their own configuration files.

For safety's sake, minds may draft changes to their own harnesses; they may never apply them. Ox proposed, Terra disposed, and neither could do the other's job. That policy is most of AI safety, practiced at the kitchen table without ceremony. Sol's four hours were metered - rented frontier cognition, billed by the token, buying something permanent: a 20% improvement capitalised into a machine I own, compounding every night it runs, at no further cost. You don't rent intelligence to do the work anymore - you rent it, briefly, to make your owned intelligence better at doing the work.

With all of that up and running, I put a question to the nightshift - one I had posed earlier that day in the closing lines of "The Emerging Cognition Surplus": What do you propose should be the next task?

It had a good long think, examined its own work, found that work wanting, and proposed, as a first step, going back to make it whole and complete. "Should I set that up as the next task?" it asked.

The whole roster in one evening: Qwen working, Sol tuning, Ox drafting, Terra applying, and me signing off on the work. Four quick minds around one slow car. The frontier models now serve as pit crew for the home stack: rent the tuner, own the tuned.

And as in racing, so here: nobody remembers the tire-changers, and no car finishes without them. But the chequered flag belongs to the slow car that runs all night.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mark pesce 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-pit-crew] indexed:0 read:3min 2026-08-25 ·