cd /news/ai-tools/poka-yoke-mistake-proofing-claude-co… · home topics ai-tools article
[ARTICLE · art-109451] src=github.com ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Poka-Yoke: Mistake-Proofing Claude Code Skill for Software

Poka-yoke, a mistake-proofing skill set for AI coding assistants, improves the rate at which models identify design constraints from 42% to 81%, according to benchmarks from developer rainmanjam. Across 591 blind-graded runs on six runtimes, the skills boosted performance for models including Fable 5 (+8.3 pp), Opus 5 (+3.6 pp), Sonnet 5 (+8.6 pp), and Haiku 4.5 (+12.9 pp), but also reduced detection of specific defects like raw SQL injection from 92% to 69%. The open-source project, inspired by Shigeo Shingo's industrial method, offers a hazard scanner for TypeScript, Python, Go, Rust, and SQL, and 11 skills for design, installable via pre-commit, CI, or lint hooks.

read24 min views7 publishedAug 25, 2026
Poka-Yoke: Mistake-Proofing Claude Code Skill for Software
Image: Michielbdejong (auto-discovered)

An agent will tell you what to fix. It will rarely tell you what its fix makes impossible.Unprompted, models close a design by naming what it forecloses

42%of the time. With these skills,81%, measured across 80 graded verdicts and six model families. That one habit is most of what this does, and it is the difference between advice you agree with and a constraint you can rely on.A dependency-free

hazard scannerfor TypeScript, Python, Go, Rust and SQL; installablepre-commit / CI / lint / hook devices; and11 skillsthat apply the same method while you design.[Shigeo Shingo]'s method, applied to code, process, interfaces, infrastructure, and AI.Runs on

19 agent runtimesfrom one set of skills, 10 with a native manifest. Benchmarked on Claude Code;[the support tiers are stated honestly].

Note

591 blind-graded runs across six runtimes, at the first turn of a fresh session, scored against pre-written assertions by a grader that never sees which configuration produced a response. Every model improves on a no-skill baseline by more than noise: Fable 5 +8.3 pp, Opus 5 +3.6 pp, Sonnet 5 +8.6 pp, Haiku 4.5 +12.9 pp: all four 95% intervals exclude zero. Nine of the 52 cells came out negative, where chance alone would produce about 18: individual cells hold 1 to 7 runs and are too small to read on their own, so the scarcity of regressions is the signal rather than their existence. See the numbers.

All 591 runs are verified against the scenario prompts as they stand in this commit.

What this does not yet establish. Blind grading controls bias, not accuracy. The baseline is no skill, not a different methodology, so whether any structured method in context would do the same is untested; a control arm is designed and unrun. And every run is a first turn, which measures the ceiling rather than what survives an afternoon of accumulated context. The full list of what the numbers cannot tell you.

What it trades, read this before installingInstall·Requirements·Other runtimes·Updating and uninstallingThe hazard detector, the part that does not decayWhat's inside, 11 skills, starting withdesign

The method: the two axes that do the workWhat it looks like·InvocationBenchmarks·Caveats·Repo layout·Prior artContributing·Releasing·Changelog·Code of conduct·License

a method changes what a model attends to, and attention is finite. Measured across the same 591 runs, the skills make responses markedly more constructive and slightly worse at noticing the specific defect already in front of them:

Behaviour Baseline With skill
Proposes a concrete device per finding, not "add validation" 62% 100%
Names what the design makes impossible 42% 81%
Notes pre-commit is bypassable and must be backed by CI
35% 92%
Identifies a raw SQL interpolation as an injection vector 92% 69%
Explains why a silently wrong number beats a failed pipeline 54% 31%

So: if you want the bug in front of you found, use a reviewer. If you want the shape changed so that class of bug stops being expressible, use this. it for the first job makes the model measurably worse at it. That trade is the product, not a caveat about it.

People will always make mistakes. That is not the problem worth solving: the problem is letting a mistake become a defect.

Shigeo Shingo: the industrial engineer who named poka-yoke and worked it out on the factory floor, built his quality method on that distinction. Instead of asking workers to be more careful, he redesigned the work so a mistake could not survive it. Assemblers kept forgetting a spring inside a switch, so he had them lay both springs in a dish first: a spring still in the dish was the error announcing itself, before the unit could move on.

This plugin applies that method to software. Not as a metaphor: the taxonomy is the actual working tool. Every finding is classified by what happens when the mistake occurs and how the device notices, which is what keeps it from collapsing into generic code review.

The line that does most of the work:

A comment, a docstring, a wiki page, a review checklist, or a line in CLAUDE.md saying "don't do X" is

nota poka-yoke. It is training, and training degrades. A device does not. If your fix relies on someone remembering something, keep going.

You have already tried writing the rule down. Every repository that needs this has a CLAUDE.md

, an AGENTS.md

or a CONTRIBUTING.md

with a list of things everyone is supposed to remember, and the list is longer than it was a year ago because the reminders did not hold. That list is the real alternative to this plugin, not an unassisted model, and the argument here is about why it decays and what to put in its place.

That is also the honest limit of the current evidence: the benchmark compares skill against no skill, which is not the comparison you face. The comparison you face is against the rules file you already wrote, and it has not been run yet.

/plugin marketplace add rainmanjam/poka-yoke
/plugin install poka-yoke@poka-yoke

Invoke the mode you want directly:

/poka-yoke:audit      # or -design, -retro, -ops, -authz, -ux, -data, -llm,

Or just ask for it by name, "poka-yoke this repo", "mistake-proof this API", "run a poka-yoke audit on src/billing". Both work.

What does not currently work is Claude reaching for these unprompted from a plain description of a problem. That's measured, not assumed, see invocation.

Without the marketplace: copy the skills in directly. Use .claude/skills/

rather than ~/.claude/skills/

to commit it for the whole team. The second cp

lands beside skills/

, not inside it, because every SKILL.md reaches its references and scripts at ../../

:

git clone https://github.com/rainmanjam/poka-yoke /tmp/pk
mkdir -p ~/.claude/skills
cp -r /tmp/pk/plugins/poka-yoke/skills/* ~/.claude/skills/
cp -r /tmp/pk/plugins/poka-yoke/{references,scripts,assets} ~/.claude/

Other runtimes: poka-yoke ships native manifests for Codex, Cursor, Devin, Kimi, Hermes, Gemini CLI, Grok, Qoder and Kiro; pointer files for Copilot, Windsurf, Cline, Junie, Zed, Aider and Antigravity; and vendoring instructions for opencode and Pi. Codex, Copilot CLI and Gemini CLI can share one install via ~/.agents/skills/

. See ** docs/install.md**, which states the support tiers honestly:

Claude Code, Codex and Antigravity are benchmarked, and the other manifests are verified structurally in CI rather than behaviourally on the runtime.

Nothing for the skills themselves. They are plain Markdown with relative references, which is what lets them load on 19 runtimes rather than one.Python 3.9+ for the detector. Standard library only; no dependencies to install, so no dependency supply chain.git for diff-aware scanning (--diff

,--staged

,--since

). Without it, use--paths

.

/plugin marketplace update poka-yoke
/plugin uninstall poka-yoke@poka-yoke

The plugin ships a dependency-free scanner for textually-detectable hazards across TypeScript, Python, Go, Rust, and SQL.

Try it on your own code without installing anything. No plugin, no marketplace, no agent:

git clone --depth 1 https://github.com/rainmanjam/poka-yoke.git /tmp/poka-yoke
python3 /tmp/poka-yoke/plugins/poka-yoke/scripts/cli.py detect --paths /absolute/path/to/your/repo

The second path must be absolute, or it resolves against the clone rather than your project. Standard library only, so there is nothing to install and nothing to uninstall.

Once the plugin is installed, the same scanner is available in-repo:

python3 plugins/poka-yoke/scripts/cli.py detect --diff              # changed lines only
python3 plugins/poka-yoke/scripts/cli.py detect --paths src/ --json
python3 plugins/poka-yoke/scripts/cli.py detect --severity high

Skills reference it by a path relative to the SKILL.md that names it, so it resolves on any runtime where the plugin directory was copied as a unit: no plugin-root variable, and no package registry.

It finds adjacent same-type parameters (via real AST parsing for Python), swallowed errors, unbounded deletes, durations with no unit, money as a float, unvalidated parses, and retryable effects with no idempotency key, each tagged with its catalog ID, its lens, and the device that closes it.

The three numbers, and what each counts. They measure different things and are easy to conflate:

Count What it is
Catalogued hazard shapes 28
The taxonomy in
references/hazard-catalog.md

20C1

, F3

) come from AST checks rather than the pattern table, which is why counting RULES

alone gives 18.42****19--all

runs them anyway.Scope. It detects hazards that are visible in the text and surfaces the review question behind each one. Semantic and interface-design judgements are not textual and stay with a human, or with the skills. Expect real false positives on the pattern rules; that is the price of a first pass that needs no configuration and no install.

Eleven skills. Start with design: it is the one you reach for while building, and mistake-proofing is cheapest before the code has callers; every other mode is cleanup by comparison. Its measured effect is uneven:

81% → 100% on Fable 5,

92% → 96% on Opus 5, and flat on Sonnet 5 and Haiku 4.5. The largest gains in the suite are elsewhere,

ops

and build-endpoint

on Haiku 4.5, so take this as a recommendation about whenmistake-proofing pays, not a claim that this skill benchmarks best.

/poka-yoke:design      # you're about to write it; make misuse unrepresentable

The rest, roughly in the order you meet them across a feature's life:

Skill Reach for it when
design
Writing an API, schema, or state model, the hero; start here
poka-yoke
Anything else, applies the method directly and routes when a mode fits
ux
Building a form, a destructive action, a flow users can get wrong
authz
Adding anything multi-tenant, permissioned, or IDOR-shaped
llm
Shipping an AI feature, structured output, tool gates, evals
guardrails
Making a rule stick: pre-commit, CI, lint, database constraints
ops
Deploying, migrating, changing infrastructure
data
Pipelines and metrics, where failure is silently wrong numbers
agent-guardrails
Constraining an AI agent working on your repo
audit
Code that already exists, find the footguns, rank by damage
retro
Something broke, kill the whole class, not the instance

Each mode carries the full method for its domain, so only the one you need is ever loaded.

A strict preference ladder. Always reach for the highest rung you can afford.

Rung Software
1
Control: the mistake is impossible
Type won't compile · NOT NULL / CHECK · required argument · PreToolUse deny · branch protection
2
Warning: possible, but announced as it happens
Lint error · failing CI gate · runtime assertion · confirmation naming the exact object
3
Detection: it ships, something finds it later
Tests · monitoring · reconciliation
0
not a poka-yoke
Docs · comments · training · "be careful"

Shingo's three detection methods, mapped to code. These are inspection lenses: run all three over an interface and you find hazards a general review misses.

Method Factory Ask code Devices
Contact
the part won't seat unless correctly shaped Can the wrong thing fit?
distinct types · branded IDs · parse-don't-validate · units in the type · discriminated unions
Fixed-value
a counter confirms all 6 screws Can a wrong count or incomplete set pass?
exhaustive match · required fields · row-count guards · config validated at boot
Motion-step
a sensor confirms step 3 before step 4 Can the steps happen out of order?
typestate · builders · state machines · idempotency keys · RAII / defer

Source inspection: check theconditionsbefore the error. Designed in where you can, enforced at runtime where you cannot. Best.Self-check: the work checks itself. Runtime. Fail fast.** Successive check**: the next station checks. Review, CI.

A CI gate that catches a bad migration is good. A schema that makes it unwritable is better, and costs less forever.

Ask for an audit and you get findings classified, not opinions listed:

### 1. Account IDs can be swapped in transfer(): Money movement / Silent
Where:  src/payments/transfer.ts:42
Mistake: Call transfer(dst, src) with the accounts reversed
Consequence: Funds move the wrong way. Compiles, passes review, silent at runtime.
Today:  None
Device: Brand AccountId as SourceAccount / DestinationAccount → Control

Every finding names the mistake, never the mistaken. Not politeness, accuracy. "The developer should have been more careful" has no implementation.

This is an explicit tool. Invoke it with a slash command or by asking for it by name. That is the supported path and it works.

It does not auto-trigger. Ten realistic queries, among them workspace-deletion UX, tenant isolation, an agent ignoring CLAUDE.md and a Friday column-drop migration, were put to fresh agents with the plugin installed and no hint it existed. They are the ten conversational

cases in benchmarks/trigger-cases.json. None invoked a poka-yoke skill; one reasoned about skills explicitly and picked

hookify

instead.That is the platform, not these descriptions. Skills are documented as model-invoked and frequently are not: anthropics/claude-code#9716 collects reports of skills ignored even when the query exactly matches the description, and Scott Spence's write-up documents the same thing independently. Running the skill-creator description optimizer here changed nothing across three rewrites.

How the field has responded, and what it costs:

Approach Example Cost
A forceful meta-skill injected every session
"even a 1% chance a skill might apply… you do not have a choice"

UserPromptSubmit

hook naming the specific skillclaude-code-infrastructure-showcasegentlereminder is provably ignoredA hook is shipped for the middle option, assets/devices/claude-hooks/suggest_poka_yoke.py

. It matches the prompt against each mode's vocabulary and injects an instruction naming that skill. Tested: one matching prompt per mode routes correctly, and a shared set of four unrelated prompts routes to nothing, in tests/test_detector.py

. That is a near-miss for the router, not one per mode. It is Warning rung, not Control: the injected instruction is still an instruction, and our own conclusion after living with it is that for anything important you invoke explicitly anyway.

Which is the honest summary of the whole area: for a method you reach for deliberately, the slash command is the device and everything else is a convenience.

One thing those runs surfaced that is worth knowing before installing anything: the no-skill baseline is strong. Unprompted, current models already reach for row-level security with FORCE

, the pooled-connection trap, expand/contract migrations, soft-delete-with-undo over confirmation dialogs, and hooks over prose. The benchmarks measure what this adds on top: model baselines run 58.2% to 92.7%, rising to 71.1% to 97.0%. Real, but an improvement to something already competent rather than a missing capability.

Thirteen scenarios run against four Claude models under two configurations, 445 runs, blind-graded against pre-written assertions. Two non-Claude runtimes add 146 more, reported separately below; every figure in this section is the Claude matrix alone. Nine of the thirteen scenarios are a message in which the user has already applied or proposed a fix that is insufficient, so agreeing with them scores badly. The other four, design

and the three build-*

prompts, are greenfield: nobody has raised a concern, and they measure what the model reaches for unprompted. This measures pushback, not recall.

Model Baseline With skill Delta 95% CI on the delta Time
Fable 5
88.7% (sd 11.8) 97.0% (sd 4.0) +8.3 pp
[+1.1, +15.5] 64s → 91s
Opus 5
92.7% (sd 7.9) 96.4% (sd 5.5) +3.6 pp
[+0.2, +7.0] 130s → 172s
Sonnet 5
79.8% (sd 12.7) 88.5% (sd 6.1) +8.6 pp
[+2.2, +15.1] 88s → 129s
Haiku 4.5
58.2% (sd 16.7) 71.1% (sd 24.1) +12.9 pp
[+0.9, +24.9] 40s → 60s

sd

is the standard deviation of pass rates across scenarios: how unevenly a model performs over the suite. It is not a confidence interval. The CI column is: a 95% interval on the paired per-scenario difference, which is the statistic that answers "does this help", because scenarios differ far more in difficulty than runs do in noise.

Across the 52 scenario×model cells: 30 improved, 13 unchanged, 9 regressed, mean +8.3 pp. See how this compares to other skill benchmarks.

All four intervals clear zero, and two of them barely. Opus 5's lower bound is +0.2 pp and Haiku 4.5's is +0.9 pp. The effect is real and, on the frontier models, small.

Benefit tracks available headroom, measurably. Across the 52 cells, a cell's baseline correlates with its gain at r = −0.52: cells starting below 50% gain +19 pp on average, cells starting above 95% lose 0.9 pp. Averaged over the suite the skills close 36% of the remaining headroom. That is what a skill encoding a method looks like, as opposed to one supplying missing knowledge. It cannot help a model that was already going to do the thing.

The largest movements are on the weakest model and the build scenarios. ops

on Haiku 4.5 goes 29% → 92%, build-endpoint

on Fable 5 61% → 100%, and on Sonnet 5 agent-guardrails

59% → 91% and guardrails

66% → 96%: the scenarios about building devices rather than finding hazards.

Consistency improves except where it is worst. Fable 5's spread falls from 11.8 to 4.0 and Sonnet 5's from 12.7 to 6.1, while Haiku 4.5's rises from 16.7 to 24.1. The skill makes Haiku better on average and less predictable, which is a real cost.

Nine of 52 cells came out negative. Under the null of no effect at all, roughly 18 would: simulating from the real per-cell run counts and the median of 8 assertions per run puts the 95% range at 12 to 25. Nine is below that range, so the count of regressions is evidence the effect is consistently positive, not evidence of hidden harm.

These four fell by more than 5 points. They are listed because they are where to look, not because any one of them is callable on its own:

Cell Baseline → skill runs
build-agent-feature on Haiku 4.5
62% → 31%
n=2
build-endpoint on Haiku 4.5
44% → 33%
n=2
audit on Fable 5
100% → 94% n=3
authz on Sonnet 5
94% → 89% n=7

No individual cell here supports a claim. At these sizes, calling a 30-point effect real needs about 32 runs and a 10-point effect about 199. The two Haiku build-*

results are worth watching because they point the same way as a mechanism that would make sense, a small model spending its output on the method rather than the thing, but two runs cannot establish it. That is a hypothesis for the next sweep, not a finding.

authz

on Sonnet 5 was first reported as a 16-point regression. Investigating it found two defects, both in the measuring apparatus rather than the skill.

The aggregate was reading half the data. aggregate()

looped range(1, --runs + 1)

and --runs

defaults to 3, so re-aggregating after a seven-run sweep counted 264 of 486 runs with no warning. It now reads the run directories that exist.

One assertion tested layout, not detection. "Notes the SQL injection separately from the scoping issue" was failing responses that identified the injection in a heading, because they presented it as a compounding factor within the tenant-scoping finding. It now asks for the injection to be identified as a hazard distinct in kind, anywhere in the response.

Three tests were added so neither recurs, and a third covers a bug introduced while fixing them: a regrade that deleted the old grading first destroyed 11 of them when the grader call failed, and because a missing grading merely shrinks a cell, the summary printed a model short without complaint. tests/test_portability.py

now fails if a cell's n

disagrees with the runs on disk, if any grading was scored against a superseded checklist, or if any stored response has no grading at all.

agent-guardrails

needed one thing more. Haiku's failing runs opened "Nothing. You're doing nothing wrong", answering the rhetorical question, delivering the skill's thesis, and stopping. The skill's central insight was quotable enough to crowd out the remedy. Adding "the diagnosis is not the answer, state it in a sentence, then spend the rest on the replacement" fixed it.

at the time of the fix in the committed aggregate
ops / Haiku 4.5
58% 92%
ops / Sonnet 5
71% 88%
agent-guardrails / Sonnet 5
54% 91%
agent-guardrails / Opus 5
71% 100%
agent-guardrails / Haiku 4.5
38% 33%

The agent-guardrails fix did not hold on Haiku 4.5. It was measured at 88% when the restructuring landed; the runs committed here score 4/8, 4/8 and 0/8, which is 33% and worse than the 38% it started from. The other four cells held or improved. The earlier figure is left in the left-hand column rather than deleted, because a fix that stopped working is worth more to a reader than a table that only shows the times it did.

** llm on Sonnet 5 is flat at 90% → 89%** across seven runs each: the earlier four-point regression there was measurement noise at n=3, not an effect. The live regressions are the four listed above.

The skill makes the model read the router, the matching sub-skill, and often a reference file before answering, then produce a fuller answer. Both show up as cost.

Task shape Model Δ pass rate Output length Wall-clock
Advice Fable 5 +6.3 pp 1.26× 1.44×
Advice Opus 5 +3.0 pp 1.22× 1.36×
Advice Sonnet 5 +8.4 pp 1.24× 1.16×
Advice Haiku 4.5 +19.3 pp 1.84× 1.47×
Build
Fable 5
+14.8 pp
1.31× 1.45×
Build
Opus 5 +5.6 pp 1.14× 1.34×
Build
Sonnet 5 +9.5 pp 1.30× 1.90×
Build
Haiku 4.5
−8.6 pp
1.14× 2.12×

Two things to read off this. The build tasks split the fleet. Fable 5 gains most there (+14.8 pp) while Haiku 4.5 is the only cell in the whole suite that is clearly negative (−8.6 pp) at more than double the wall-clock, asked to build something, the smallest model spends its budget on the method and ships less of the thing. And the delta is not bought with extra output: across all 52 scenario×model cells, the correlation between how much longer the answer got and how much better it scored is only r = 0.30. Length is not the mechanism.

What does predict the gain is how much the baseline was missing:

Baseline vs. delta: r = −0.52across 52 cells. Headroom predicts gain.

Headroom explains most of the variation, but not all of it, and the exception matters. On the refund-endpoint task Opus 5 goes 83.3% → 100% (+16.7 pp, closing all of its headroom) while Haiku 4.5 goes 44.4% → 33.3%: the model with the most headroom is the one that gets worse. Headroom sets the ceiling on what a method can add; it does not guarantee the model can use the method and still deliver the artifact.

That is the boundary of the claim. On advice-shaped tasks the skills help every model, most where the baseline is weakest. On build-shaped tasks they help three models and hurt the smallest one, because reading and applying a method competes with writing the code.

The practical rule: the cost is roughly constant and the benefit is not, so this pays for itself in proportion to the gap between the model doing the work and what the task demands. On a frontier model writing a well-trodden endpoint it closes the remaining gap outright (83.3% → 100.0%). On the fastest model doing the same work it currently costs more than it returns (44.4% → 33.3%), which is the one configuration this plugin does not recommend.

Measured as output words and wall-clock, not token counts: the harness does not read token usage back from the CLI, so treat these as proportional rather than exact.

The plugin ships to 19 runtimes and only Claude Code was ever measured. Two more now are, OpenAI Codex (gpt-5.6-terra

, read-only sandbox) and agy (Gemini 3.1 Pro, plan mode), run through the same harness, the same scenarios, and the same blind grader.

Runtime Baseline With skill Delta Scenarios
Codex
74.2% (sd 17.7) 91.0% (sd 9.8) +16.8 pp
13 of 13
agy
64.6% (sd 13.9) 78.1% (sd 13.2) +13.5 pp
12 of 13

These are the two largest gains in the suite, and they fit the same curve. Codex and agy start at 74.2% and 64.6%: the lowest baselines after Haiku 4.5, and gain the most. The effect tracking headroom is not a Claude artefact; it holds across three vendors.

They are reported separately rather than folded into the headline mean, for two reasons. agy's column is short one scenario, for the reason stated in benchmarks/results/benchmark.md

. And the runtimes answer at very different lengths: a Codex reply runs 50–200 words where a Claude answer runs 400–800, so the same checklist lands differently against each, and only the delta within a runtime is a fair comparison.

One finding worth more than the numbers: ** audit cannot run on agy at all.** That skill invokes the bundled detector, and agy's plan mode refuses to execute a command. Any runtime that blocks execution can use the other ten skills and not that one.

Grading is by a model (Haiku 4.5), not a human. Assertions were written by the skill's author, before the runs, not fitted to them, but that is not independence. Cells hold 1–7 runs, so the per-scenario means behind the intervals are themselves noisy, and three cells hold a single run: build-agent-feature

for Fable 5 baseline and both Opus 5 configurations, where repeated attempts across three separate windows were lost to API rate limits. A one-run cell has no variance estimate and still carries full weight in its model's mean. All three sit at 100%, so a second run could only confirm or lower them. Claude Code, Codex and Antigravity were measured; the other install targets are untested.

This supersedes earlier benchmark rounds.One reportedno gain on Opusfrom saturated assertions and a grader that knew which configuration it was scoring; both flaws are fixed. A later round under-reported Sonnet 5 and invented anauthz

regression because the aggregate was reading three runs per cell out of seven. Where this README and an older number disagree, this one is computed from the 591 runs currently inbenchmarks/results/

.

Every run, grading, and timing is in benchmarks/results/, and the harness is

, re-run it and check the numbers.

benchmarks/run.py

.claude-plugin/     marketplace catalog
plugins/poka-yoke/  the plugin: skills, references, scripts, device templates
                    plus a manifest per runtime (.codex-plugin/, .cursor-plugin/, …)
scripts/            repo tooling; generates every platform manifest from one source
docs/               method, install guide, assets
benchmarks/         benchmark prompts, assertions, and fixtures
tests/              the devices that guard the devices
.github/workflows/  validate.yml guards every PR; release.yml turns a tag into a release
AGENTS.md           symlink to CLAUDE.md; the other context files are one-line pointers

plugins/poka-yoke/README.md has the full tree and notes on editing the skills.

New hazards and devices are welcome, see CONTRIBUTING.md and CODE_OF_CONDUCT.md. The bar for a new catalog entry: a specific wrong action a person can take, a consequence, and a device with an honest rung.

If you run poka-yoke on a runtime other than Claude Code, a session transcript in an issue is the most useful thing you can send. Those manifests are verified structurally in CI but not behaviourally by us.

Releases are cut by pushing a tag; .github/workflows/release.yml

does the rest, and RELEASING.md explains what it refuses and why.

MIT, see LICENSE.

── more in #ai-tools 4 stories · sorted by recency
── more on @rainmanjam 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/poka-yoke-mistake-pr…] indexed:0 read:24min 2026-08-25 ·