# pstack's Playbooks, Explained as Factorio Factories

> Source: <https://alexop.dev/posts/pstack-playbooks-as-factories/>
> Published: 2026-10-08 00:00:00+00:00

pstack's Playbooks, Explained as Factorio Factories

Published: at

If you have played Factorio, you know that every factory has a shape. A smelter line is long and straight. A circuit factory has loops. A train station has buffers. You can tell what a factory makes just by looking at it from above.

I’ve been using pstack, a plugin by poteto that turns a coding agent into something closer to a careful engineer. After a while I noticed the same thing: every pstack playbook has a shape. A bug fix has a loop back to the start. A refactoring has a gate that throws work away. A hillclimb is one big loop.

pstack is a set of skills. You don’t have to learn most of them, because one skill, /poteto-mode, runs the rest for you. You describe the task and how you’ll know it’s done. poteto-mode matches the task to one of 23 playbooks, copies the playbook’s steps into a todo list, and calls the other skills (how, architect, arena, interrogate, and so on) as each step needs them.

The idea behind it is simple: AI writes a lot of code fast, and most of it is slop. pstack doesn’t try to make the agent faster. It makes the agent prove its work.

pstack was built for Cursor. I use Claude Code, so I ported it: pstack-claude. The skills and playbooks are poteto’s. The port only swaps the Cursor-specific parts.

Every playbook below is one factory. Work flows from left to right.

BeltCarries work. Items queue when the machine ahead is busy.

InserterMoves an item from a belt into a machine, and out again.

AssemblerA step that builds: code, a fix, a candidate.

LabA step that checks: repro, verify, measure.

Drafting tableA step that plans: architect, or a brief for a worker.

RadarReads the code (how).

RecyclerDeletes code: subtract before you add.

Chemical plantAdversarial review (interrogate).

Ghost blueprintNever ships: throwaway code, or a PR that won't happen.

SplitterSends items whose dot matches its own onto the side belt.

Red loopRework: the item goes back around for another try.

Rocket siloA merged PR.

ChestOutput that isn't code, like a report.

Dots and numbersA dot is a verdict or a variant: green pass, red fail, orange contested. A number is a count: suspects, ms, PR #.

The caption under each map says what to watch. The list under it is the real playbook: click a step and the factory pauses and lights up the machines that do it. Click it again to resume. The numbers are made up. The shapes are not: each one follows the steps in poteto’s playbook files.

This is the part you actually type. You never pick a playbook yourself. The same command with different words lands in a different factory:

What you type after /pstack:poteto-mode

Playbook

users get two notifications after a retry. repro first, then fix and verify.

Bug fix

investigate why background jobs time out every few hours. don’t change any code yet.

Investigation

move parsing into one module, zero behavior change.

Refactoring

startup takes 1.8s on this fixture. trace it, show me before and after.

Perf

In Cursor the command is /poteto-mode. In Claude Code, with my port, it’s /pstack:poteto-mode.

Five shapes of task get their own splitter here. The real list has 23 playbooks. When none of them fits, the task doesn’t default to Feature. It falls through to figure-it-out, a skill that designs a bespoke playbook for that one task.

A read-only question (“why does X happen”, “don’t change any code”) is what routes it to Investigation.

There is no silo, only a blueprint marked “No PR”. I like that this is a separate playbook, because “explain this to me” and “change this” are different jobs. An agent that mixes them up starts editing files you only asked about. When the answer turns out to need a code change, the playbook hands it back to be re-routed to Bug fix or Feature.

A reported defect, plus “repro first”, routes it to Bug fix.

Nothing gets past the red lab until the bug actually fires. If it won’t reproduce, the agent adds logging until it does. It doesn’t ask you to reproduce it.

The long green belt is the point. The fix only counts when the original repro passes on the same surface. A unit test that passes somewhere else doesn’t count.

The architect only runs when the fix crosses a function boundary. A one-line fix inside one function skips it.

Last detail: the commit machine drops two items. First a red test, then the fix. In git history the failing test lands before the fix, so anyone can check out the commit before and watch it fail.

When there are several valid ways to build something, pstack doesn’t let one agent pick. It runs an arena: the same brief goes to three runners, each surfaces its own shape, and a judge takes one as the base and grafts the best parts of the other two into it. When there’s one obvious shape, the arena is skipped.

Designs that are contested (orange dot) take a detour through interrogate, where reviewers try to break the change. More on both skills below.

“Zero behavior change” is what routes it to Refactoring.

A refactoring must not change what the code does, so the first machine is a pin: a test that records what the old module does today. Two things stand out:

The recycler comes before the assembler. pstack deletes dead code, one-caller wrappers and old validators before it builds the new shape. Subtract before you add.

There are two ways to fail. A broken pin goes back for another small step. A change that keeps the output but doesn’t read easier gets reverted. A refactoring that doesn’t lower reader load has no reason to exist.

A perf fix starts with a baseline, and that number gets vetted before anyone trusts it (benchmark-checklist). Then the slow path tries the seven performance mantras in order, cheapest first:

Don’t do it.

Do it, but don’t do it again.

Do it less.

Do it later.

Do it when they’re not looking.

Do it concurrently.

Do it cheaper.

What you type:

```
/pstack:poteto-mode startup takes 1.8s on this fixture. trace it, fix the measured cause, show me before and after.
```

A measured slowness is what routes it to Perf.

The playbook has a stop rule: when an earlier mantra meets the target, stop. That’s the bypass belt. Nobody reaches for “do it concurrently” when “don’t do it” already solved the problem. Then the green lab measures again, and the before and after go into the PR. No number, no PR.

A metric, a target and a floor on attempts is the shape Hillclimb asks for.

Hillclimb is for pushing one metric up over many attempts. The stop rule has two halves on purpose. Without the floor on attempts, a lucky first try ends the run. Every lap writes one row to decision.tsv, so you can read the whole run the next morning. A win only counts if the regression tests stay green too.

After three losses in a row, watch for “plateau → pivot category”. That’s the playbook telling the agent not to stop at the first plateau, but to try a different kind of idea.

“Prototype” is one of the words that routes straight to this playbook.

A prototype exists to answer one question, like “which layout?”. No question, no prototype. The variants are throwaway code in a scratch folder: no framework, no tests. All three sit behind one switcher, so the agent can screenshot and compare them. The output is a decision, not code.

Shipping is for a stack of PRs that depend on each other. Every PR gets its own verifier: an agent that didn’t write the code. A green CI run doesn’t count as a verdict. The ceiling is the rule I find most useful: a verified PR on top of an unverified one isn’t safe to land, however green it looks.

A project you hand over for days (“own it until…”) is what routes it to Orchestrate.

The biggest playbook. One coordinator chat runs a project that takes days and many PRs. It starts with a goal you can count, like “16 units merged”. The coordinator never writes code. Its product is the brief, because a worker can’t ask it a question.

The pilot tests the brief and the verify recipe while a mistake costs one agent instead of fifty. After that, a rolling window beats blocking batches: a batch waits for its slowest worker, a window refills each worker the moment it’s done. Landing runs the whole time. It’s never a final phase.

This is what’s inside the chemical plant in the Feature factory. One reviewer per configured model reads the same diff with the same prompt and rubric. The adversarial signal comes from the reviewers being different, not from personas. A finding two reviewers raise on their own is the strongest signal. The lead then sorts every finding into four bins, and nothing gets applied automatically.

One honest caveat: upstream pairs a Claude model with a Grok model. In my Claude Code port, all subagents run on Claude models, so the reviewers differ by separate context, not by model family.

This is the fan of A/B/C builders in the Feature factory. The prompt is the contract, and a short rubric that only the picker sees decides the base. The base is the candidate a future maintainer can extend most easily. If all candidates converge, that’s agreement, and nothing gets grafted. If they wildly diverge, the frame was too vague, so the task goes back to be re-framed.

This is the drafting table in the Bug fix and Feature factories. It designs before code: at least two structurally different designs, screened against a list of red flags, then one sketch to implement against. When implementation keeps hitting the same workaround, the sketch gets thrown out instead of patched.

Swarm fans out workers, one brief each, and returns one report. A result without its evidence doesn’t count: that worker gets respawned once, and after a second miss the slice is a gap. A gap is not a pass.

Listing callers is not the job, the agent can grep those in a second. Blast-radius looks for the one fact a change is safe because of, and then proves it by running real code. The staircase is the skill’s own confidence ladder. A writeup that sounds right is worth nothing until the fact reaches step 4.

Look at them again. Every factory has a machine whose only job is to check, and every checker can send work back or stop the line:

Bug fix: the repro lab.

Feature: the judge and interrogate.

Refactoring: the pin gate.

Perf and hillclimb: the measuring labs.

Shipping: the verifiers and the ceiling.

None of the playbooks make the assembler faster. The coding part was never the slow part. In my Delivery Factory, AI made dev fast and the work piled up in review and QA. pstack puts the review and QA inside the agent’s own factory, so less broken work reaches you.
