{"slug": "pstack-s-playbooks-explained-as-factorio-factories", "title": "pstack's Playbooks, Explained as Factorio Factories", "summary": "Pstack, a Cursor plugin by poteto that routes coding-agent tasks through 23 playbooks, was ported to Claude Code as pstack-claude, swapping only the Cursor-specific parts while keeping poteto's skills and playbooks. The tool's /poteto-mode skill matches a described task to one of the 23 playbooks, copies its steps into a todo list, and calls supporting skills such as how, architect, arena and interrogate as each step needs them, with unmatched tasks falling through to a figure-it-out skill that designs a bespoke playbook. The port's author frames the design goal as making the agent prove its work rather than write code faster.", "body_md": "pstack's Playbooks, Explained as Factorio Factories\n\nPublished: at\n\nIf you have played Factorio, you know that every factory has a shape. A smelter line is long and straight. A circuit factory has loops. A train station has buffers. You can tell what a factory makes just by looking at it from above.\n\nI’ve been using pstack, a plugin by poteto that turns a coding agent into something closer to a careful engineer. After a while I noticed the same thing: every pstack playbook has a shape. A bug fix has a loop back to the start. A refactoring has a gate that throws work away. A hillclimb is one big loop.\n\npstack is a set of skills. You don’t have to learn most of them, because one skill, /poteto-mode, runs the rest for you. You describe the task and how you’ll know it’s done. poteto-mode matches the task to one of 23 playbooks, copies the playbook’s steps into a todo list, and calls the other skills (how, architect, arena, interrogate, and so on) as each step needs them.\n\nThe idea behind it is simple: AI writes a lot of code fast, and most of it is slop. pstack doesn’t try to make the agent faster. It makes the agent prove its work.\n\npstack was built for Cursor. I use Claude Code, so I ported it: pstack-claude. The skills and playbooks are poteto’s. The port only swaps the Cursor-specific parts.\n\nEvery playbook below is one factory. Work flows from left to right.\n\nBeltCarries work. Items queue when the machine ahead is busy.\n\nInserterMoves an item from a belt into a machine, and out again.\n\nAssemblerA step that builds: code, a fix, a candidate.\n\nLabA step that checks: repro, verify, measure.\n\nDrafting tableA step that plans: architect, or a brief for a worker.\n\nRadarReads the code (how).\n\nRecyclerDeletes code: subtract before you add.\n\nChemical plantAdversarial review (interrogate).\n\nGhost blueprintNever ships: throwaway code, or a PR that won't happen.\n\nSplitterSends items whose dot matches its own onto the side belt.\n\nRed loopRework: the item goes back around for another try.\n\nRocket siloA merged PR.\n\nChestOutput that isn't code, like a report.\n\nDots and numbersA dot is a verdict or a variant: green pass, red fail, orange contested. A number is a count: suspects, ms, PR #.\n\nThe caption under each map says what to watch. The list under it is the real playbook: click a step and the factory pauses and lights up the machines that do it. Click it again to resume. The numbers are made up. The shapes are not: each one follows the steps in poteto’s playbook files.\n\nThis is the part you actually type. You never pick a playbook yourself. The same command with different words lands in a different factory:\n\nWhat you type after /pstack:poteto-mode\n\nPlaybook\n\nusers get two notifications after a retry. repro first, then fix and verify.\n\nBug fix\n\ninvestigate why background jobs time out every few hours. don’t change any code yet.\n\nInvestigation\n\nmove parsing into one module, zero behavior change.\n\nRefactoring\n\nstartup takes 1.8s on this fixture. trace it, show me before and after.\n\nPerf\n\nIn Cursor the command is /poteto-mode. In Claude Code, with my port, it’s /pstack:poteto-mode.\n\nFive shapes of task get their own splitter here. The real list has 23 playbooks. When none of them fits, the task doesn’t default to Feature. It falls through to figure-it-out, a skill that designs a bespoke playbook for that one task.\n\nA read-only question (“why does X happen”, “don’t change any code”) is what routes it to Investigation.\n\nThere is no silo, only a blueprint marked “No PR”. I like that this is a separate playbook, because “explain this to me” and “change this” are different jobs. An agent that mixes them up starts editing files you only asked about. When the answer turns out to need a code change, the playbook hands it back to be re-routed to Bug fix or Feature.\n\nA reported defect, plus “repro first”, routes it to Bug fix.\n\nNothing gets past the red lab until the bug actually fires. If it won’t reproduce, the agent adds logging until it does. It doesn’t ask you to reproduce it.\n\nThe long green belt is the point. The fix only counts when the original repro passes on the same surface. A unit test that passes somewhere else doesn’t count.\n\nThe architect only runs when the fix crosses a function boundary. A one-line fix inside one function skips it.\n\nLast detail: the commit machine drops two items. First a red test, then the fix. In git history the failing test lands before the fix, so anyone can check out the commit before and watch it fail.\n\nWhen there are several valid ways to build something, pstack doesn’t let one agent pick. It runs an arena: the same brief goes to three runners, each surfaces its own shape, and a judge takes one as the base and grafts the best parts of the other two into it. When there’s one obvious shape, the arena is skipped.\n\nDesigns that are contested (orange dot) take a detour through interrogate, where reviewers try to break the change. More on both skills below.\n\n“Zero behavior change” is what routes it to Refactoring.\n\nA refactoring must not change what the code does, so the first machine is a pin: a test that records what the old module does today. Two things stand out:\n\nThe recycler comes before the assembler. pstack deletes dead code, one-caller wrappers and old validators before it builds the new shape. Subtract before you add.\n\nThere are two ways to fail. A broken pin goes back for another small step. A change that keeps the output but doesn’t read easier gets reverted. A refactoring that doesn’t lower reader load has no reason to exist.\n\nA perf fix starts with a baseline, and that number gets vetted before anyone trusts it (benchmark-checklist). Then the slow path tries the seven performance mantras in order, cheapest first:\n\nDon’t do it.\n\nDo it, but don’t do it again.\n\nDo it less.\n\nDo it later.\n\nDo it when they’re not looking.\n\nDo it concurrently.\n\nDo it cheaper.\n\nWhat you type:\n\n```\n/pstack:poteto-mode startup takes 1.8s on this fixture. trace it, fix the measured cause, show me before and after.\n```\n\nA measured slowness is what routes it to Perf.\n\nThe playbook has a stop rule: when an earlier mantra meets the target, stop. That’s the bypass belt. Nobody reaches for “do it concurrently” when “don’t do it” already solved the problem. Then the green lab measures again, and the before and after go into the PR. No number, no PR.\n\nA metric, a target and a floor on attempts is the shape Hillclimb asks for.\n\nHillclimb is for pushing one metric up over many attempts. The stop rule has two halves on purpose. Without the floor on attempts, a lucky first try ends the run. Every lap writes one row to decision.tsv, so you can read the whole run the next morning. A win only counts if the regression tests stay green too.\n\nAfter three losses in a row, watch for “plateau → pivot category”. That’s the playbook telling the agent not to stop at the first plateau, but to try a different kind of idea.\n\n“Prototype” is one of the words that routes straight to this playbook.\n\nA prototype exists to answer one question, like “which layout?”. No question, no prototype. The variants are throwaway code in a scratch folder: no framework, no tests. All three sit behind one switcher, so the agent can screenshot and compare them. The output is a decision, not code.\n\nShipping is for a stack of PRs that depend on each other. Every PR gets its own verifier: an agent that didn’t write the code. A green CI run doesn’t count as a verdict. The ceiling is the rule I find most useful: a verified PR on top of an unverified one isn’t safe to land, however green it looks.\n\nA project you hand over for days (“own it until…”) is what routes it to Orchestrate.\n\nThe biggest playbook. One coordinator chat runs a project that takes days and many PRs. It starts with a goal you can count, like “16 units merged”. The coordinator never writes code. Its product is the brief, because a worker can’t ask it a question.\n\nThe pilot tests the brief and the verify recipe while a mistake costs one agent instead of fifty. After that, a rolling window beats blocking batches: a batch waits for its slowest worker, a window refills each worker the moment it’s done. Landing runs the whole time. It’s never a final phase.\n\nThis is what’s inside the chemical plant in the Feature factory. One reviewer per configured model reads the same diff with the same prompt and rubric. The adversarial signal comes from the reviewers being different, not from personas. A finding two reviewers raise on their own is the strongest signal. The lead then sorts every finding into four bins, and nothing gets applied automatically.\n\nOne honest caveat: upstream pairs a Claude model with a Grok model. In my Claude Code port, all subagents run on Claude models, so the reviewers differ by separate context, not by model family.\n\nThis is the fan of A/B/C builders in the Feature factory. The prompt is the contract, and a short rubric that only the picker sees decides the base. The base is the candidate a future maintainer can extend most easily. If all candidates converge, that’s agreement, and nothing gets grafted. If they wildly diverge, the frame was too vague, so the task goes back to be re-framed.\n\nThis is the drafting table in the Bug fix and Feature factories. It designs before code: at least two structurally different designs, screened against a list of red flags, then one sketch to implement against. When implementation keeps hitting the same workaround, the sketch gets thrown out instead of patched.\n\nSwarm fans out workers, one brief each, and returns one report. A result without its evidence doesn’t count: that worker gets respawned once, and after a second miss the slice is a gap. A gap is not a pass.\n\nListing callers is not the job, the agent can grep those in a second. Blast-radius looks for the one fact a change is safe because of, and then proves it by running real code. The staircase is the skill’s own confidence ladder. A writeup that sounds right is worth nothing until the fact reaches step 4.\n\nLook at them again. Every factory has a machine whose only job is to check, and every checker can send work back or stop the line:\n\nBug fix: the repro lab.\n\nFeature: the judge and interrogate.\n\nRefactoring: the pin gate.\n\nPerf and hillclimb: the measuring labs.\n\nShipping: the verifiers and the ceiling.\n\nNone of the playbooks make the assembler faster. The coding part was never the slow part. In my Delivery Factory, AI made dev fast and the work piled up in review and QA. pstack puts the review and QA inside the agent’s own factory, so less broken work reaches you.", "url": "https://wpnews.pro/news/pstack-s-playbooks-explained-as-factorio-factories", "canonical_source": "https://alexop.dev/posts/pstack-playbooks-as-factories/", "published_at": "2026-10-08 00:00:00+00:00", "updated_at": "2026-10-08 19:50:34.288390+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "ai-products"], "entities": ["pstack", "poteto", "Cursor", "Claude Code", "pstack-claude", "/poteto-mode", "figure-it-out"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/pstack-s-playbooks-explained-as-factorio-factories", "markdown": "https://wpnews.pro/news/pstack-s-playbooks-explained-as-factorio-factories.md", "text": "https://wpnews.pro/news/pstack-s-playbooks-explained-as-factorio-factories.txt", "jsonld": "https://wpnews.pro/news/pstack-s-playbooks-explained-as-factorio-factories.jsonld"}}