cd /news/ai-tools/forespec-catches-what-your-ai-coder-… · home topics ai-tools article
[ARTICLE · art-131787] src=github.com ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Forespec – catches what your AI coder didn't know to ask

Forespec, an early-build tool from developer SteveWeed79, runs inside AI coding agents such as Claude Code to surface non-obvious requirements — atomic stock holds, tenant isolation, payment webhook authenticity — before code is written, then grades what was built. The tool ships five archetypes, a reasoning verifier, a plan/interrogator, a PR gate, and a greenfield on-ramp, and can be tried with no install or API key via the command "npx forespec demo"; standalone grading requires an API key, while the Claude Code plugin path runs on an existing Claude Code subscription.

read12 min views2 publishedSep 16, 2026
Forespec – catches what your AI coder didn't know to ask
Image: Michielbdejong (auto-discovered)

Forespec keeps you pointed in the right direction — at the start of a feature, along the way as it grows, and over time by remembering your past results to catch code degradation before it compounds. It's the domain foresight a senior engineer brings, codified, and kept live through your whole build.

You ship fast with AI coding tools. They build what you ask — not the non-obvious thing your kind of app actually requires: an ecommerce checkout needs an atomic stock hold; a SaaS needs tenant isolation; a payment webhook needs authenticity. Miss one and you find out in month three, doing surgery on a live flow. Forespec surfaces those requirements before you build, hands your AI coder a gotcha-aware spec, then grades what got built and tracks how each part moves run-over-run — so the foresight arrives on time, and stays live.

Forespec is not a security scanner. Security is one row of what it checks; the rest is correctness, data-modeling, reliability, and design — the whole backbone your archetype requires. It runs inside the coding agent you already use (Claude Code first), so the foresight lands while the code is being written rather than in a review three days later.

Status: early build. Five archetypes, the reasoning verifier, the plan/interrogator, the PR gate, and the greenfield on-ramp are here and runnable. New to this kind of project? SETUP.md gets you going on Windows, step by step.

No install, no API key, nothing to configure — this replays the verifier on a bundled vulnerable-checkout example so you can see exactly what a real grade looks like:

npx forespec demo

It flags the non-obvious criticals (a Stripe call with no idempotency key → double charges; a stock race; an order marked paid on the unverified client return), passes the one that's actually fine (card data never touches your server), and surfaces a required piece you haven't built yet — the discernment a grader you'd trust with "is this shippable?" has to earn. Then point it at your own repo:

The shortest path to a real grade on your own code. No API key, no metered cost — it runs on the Claude Code subscription you already have:

/plugin marketplace add SteveWeed79/forespec
/plugin install forespec@forespec
/forespec:plan add checkout     # before you build — what does this actually require?
/forespec:verify                # after you build — what's shippable, what's not

It also loads on its own the moment you start writing payment, auth, tenancy, upload, LLM, or Supabase/Firebase code, so the requirement arrives at write time — an idempotency key passed on the payment intent is an argument; retrofitted into a live payment path it's surgery. Full details, including how to drive it from any other agent: docs/claude-code-plugin.md.

This is also the better grader, not a cheap fallback. The API path packs keyword-ranked files into a character budget and grades the blob, so it can only cite a file path. An agent has grep and read: it follows the call chain into the middleware that supposedly verifies the signature, and cites file:line — a finding you can fix instead of one you have to go re-investigate.

For scripting, or when there's no agent in the loop:

forespec start "an online store with checkout"

forespec init

forespec plan "add checkout flow"

forespec verify              # standalone grading needs an API key — see below
forespec verify --html       # …drop a visual report you can open in a browser
forespec gate --help         # wire the PR/CI gate that comments on every pull request

start and plan are the foresight-before-build half: they surface the non-obvious, archetype-required properties — the "decide first" questions and acceptance criteria — and hand them to your AI coder as an ordered, dangerous-pieces-first spec. verify and gate are the keeps-it-honest half: they grade those same checkpoints on the real code, flag what's unsafe, surface the required backbone you haven't built yet, and — reading the calibration store — show how each checkpoint moved since your last run, so a regression doesn't slip through. Over time forespec proficiency reads how much judgment you've shown per domain (self-facing only) and plan adapts how much it explains — full where you're learning, terse where you're fluent.

start/ init read only metadata (dependencies, paths, schema-model names) or your one-line description — never your code — to pick the archetype.

Grading needs a reasoning verifier, and there are two ways to get one:

  • The plugin (above) — your coding agent grades the repo and hands the verdicts back through the same roll-up, gaps report and calibration store. No key, no cost, and it can citefile:line . Run it as/forespec:verify . Measured on the same corpus as the API path:0 false-greens across 152 critical-bad trials → 95% upper bound ≤2.0% , with 100% outcome agreement across two independent runs. Caveats — that measures the grading contract on snippets, not the repo-navigation advantage, and your session's model is the grader — are inVALIDATION-NOTES.md .
  • An API key — setANTHROPIC_API_KEY +ANTHROPIC_MODEL andverify calls the model directly. It carries the measured bar: 0 false-greens on 52 critical bad cases, rule-of-three 95% upper bound ≤ 2.9% (seeVALIDATION-NOTES.md ). That bar covers the ecommerce/universal corpus; the newersaas /ai-app /baas archetypes arefirst-pass validated (full rule-of-three pending).

With neither, verify refuses. It won't fake a grade — a page of keyword-matched "risky" verdicts looks like a verdict no matter how it's labelled, and the point of this tool is a number you can act on. The mock keyword baseline still exists as the dumb bar a real verifier has to beat, but you have to ask for it by name (--adapter mock), and it can never certify a merge. Full walkthrough: repo-verify/README.md.

CI runs on your subscription too. claude setup-token produces an OAuth token that authenticates with a Pro/Max/Team/Enterprise plan, so the PR gate bills against the plan you already have rather than a metered key — docs/ci-gate-agent.md. The merge decision stays deterministic: the agent produces verdicts, forespec gate reads them and decides.

What works with no verifier at all, right now, no setup: forespec demo (a real graded run on a bundled example), forespec plan "<feature>" (what the feature actually requires, before you build it), forespec init (archetype detection from metadata), and forespec checkpoints (the standard itself, as JSON). The foresight half of the product costs nothing to try.

  • Pointstart /plan interrogate the domain and emit a gotcha-aware spec (atomic hold before Stripe, data-model shape before the features built on it), most-foundational first.
  • Build — you, or your AI coder, build against that spec.
  • Verifyverify /gate grade what actually got built, flag what's unsafe, and name the required backbone you haven't reached yet.
  • Remember — every run writes to a local calibration store behind a strictpattern / instance wall , so the next run can tell you whatmoved (catching a regression) and, over time, sharpen the foresight itself.

That last step is the point: the standard isn't a static checklist — it compounds on your work (and, opt-in later, across a shared pattern pool), while your project's specifics never leave your machine.

Pointed at 8 public OSS repositories it had never seen — spanning all five archetypes, including a 4,332-file Python codebase — the plugin's grader produced 147 verdicts: 18 findings, 112 passes, 17 N/A. Every finding was then checked against the source by hand:

  • 12 of 18 findings hand-verified. 0 fabrications. Everyfile:line pointed at real code that said what the verdict claimed.
  • 0 of 147 verdicts came back without evidence. Every one citedfile:line .
  • It comes back clean on clean code. Documenso, LibreChat and supabase-js produced zero findings; four of Documenso's strongest passes were falsification-tested and held.
  • The full ledger — including four defects it found in Forespec itself — is indocs/oss-audit-2026-09.md . Reproduce any run withnode verifier-eval/repo-audit.mjs --repo <path> .

Pointed at a real production ecommerce app (a codebase it had never seen, not a fixture), forespec verify returned one blocking critical: the Stripe checkout-session creation call carried no idempotency key, so a double-click or a client retry could open two charges for one cart. Confirmed real by direct code review. It also caught a subtler one while passing that checkpoint at level 9 — a refund idempotency key with no per-refund nonce, so two equal-amount partial refunds collide and reconcile wrong.

Every finding was then independently re-checked against the actual source. The result, recorded honestly in docs/real-repo-audit-2026-07.md:

  • 0 fabricated findings, 0 false-greens — every flag pointed at real code.
  • The bias is over-severity, not fabrication — it would rather flag something you've already handled than miss something you haven't. That's the safe direction, and the candor a grader you trust with*"is this shippable?"* has to earn.

Not "perfect" — honest. That's the whole point.

File What it is
docs/ci-gate-agent.md The PR gate on a Claude subscription instead of an API key — and why the merge decision stays with pr-gate.mjs , not the model.
docs/oss-audit-2026-09.md The field report. 8 public OSS repos graded by the plugin, every finding checked by hand — including the four defects it found in Forespec itself.
docs/claude-code-plugin.md The front door. How the plugin turns your coding agent into the verifier — no API key — and how to drive the same path from any other agent.
FORESPEC-2.md The vision: the full architecture and the moat argument. Superseded on build sequence by the build order below.
forespec.buildorder-2.md The authoritative roadmap. Phases 0–7, verifier-first, each phase shippable on its own. When any doc disagrees onwhat to build in what order , this one governs.
forespec.calibration-1.md The calibration loop that turns invented weights into ones earned on real work, and the seam that lets solo data later join a shared pool without a rewrite.
library/ The shared checkpoint library — every checkpoint definition (auth, payment, data, design, ai, baas, …), authored once and reused across archetypes.resolve.mjs composes a manifest + the library into a full archetype.
archetype.ecommerce.json The ecommerce archetype manifest : 20 backbone + 7 design checkpoints from the library, each with its severity for this domain. Resolves to the durable standard a verifier grades against.
archetype.ecommerce.design.json The instrumented design layer: design checkpoints decomposed into weighted, measurable sub-signals → a computed 0–10 composite. Itsmodel_scored signals are deferred experiments until calibration earns them.
archetype.saas.json The SaaS / subscription manifest — 26 checkpoints, all but 3 reused , 3 SaaS-specific (tenant isolation, entitlement integrity, subscription lifecycle).
archetype.portfolio.json The portfolio / content manifest — 14 checkpoints, 100% composed from the shared library (design + web + the universal set), zero new authoring.
archetype.ai-app.json The AI / LLM app manifest — 12 checkpoints,5 AI-specific (prompt injection, output handling, tool-use safety, cost controls, data boundary) + 7 reused.
archetype.baas.json The Backend-as-a-Service (Supabase / Firebase) manifest — 10 checkpoints,3 BaaS-specific (RLS enforced, client trust boundary, privileged-key exposure) + 7 reused.

Reading order: FORESPEC-2.md (the why) → forespec.buildorder-2.md (the how and the order, the plan of record) → forespec.calibration-1.md (the layer that keeps every score honest over time) → library/ + archetype.ecommerce.json (the standard itself).

  • Pattern / instance wall. Transferable patterns and never-leaves-the-project instance data live in separate stores from the first write — the legal and ethical line, and the thing that makes a shared pattern pool safe to opt into later.

  • Honesty mechanic. Every score reports its level, the gap to the next, and its basis. A score that can't state its basis doesn't ship.

  • Stable, namespaced checkpoint ids (e.g.payment.webhook_authenticity ,ecommerce.checkout.atomic_stock_hold ) are permanent contracts. Archetypes areversioned ; checkpoints arenever silently renamed — that's how calibration history survives.

  • Library + manifest composition. Checkpoints are defined once inlibrary/ andcomposed per archetype (a manifest of{ ref, severity } ), so a fix to a shared checkpoint improves every archetype and a new archetype reuses instead of copies.

  • Doc naming. Prose specs areforespec.<topic>-<n>.md ;FORESPEC-2.md is the top-level vision. The trailing-1 /-2 are iteration numbers (higher = later).

  • Archetype versioning. Each manifest and library file carries a semverversion / format tag. Checkpointids are permanent contracts — bump the version when a definition changes, never rename an id.

  • $schema. Files declare a forespec/… format tag (forespec/archetype/v2 ,forespec/checkpoint-library/v1 , …) — the engine's internal contract, stable across the brand. JSON Schemas that validate them live inschemas/ .

  • bin/forespec.mjs — the unified forespec CLI (start /init /detect /plan /verify — add--html for a visual report — /design /gate /feedback /calibrate /proficiency ), exposed fornpx .

  • library/ — shared checkpoint library +resolve.mjs (compose a manifest + library into a full archetype:node library/resolve.mjs archetype.ecommerce.json ).

  • schemas/ — JSON Schema for the library, the manifest, the resolved archetype, and the instrumented design layer, plus a zero-dependency invariant validator (node schemas/validate.mjs ).

  • verifier-eval/ — the verifier-accuracy harness. A labeled good/bad fixture corpus for the critical backbone checkpoints, plus a runner that measures a verifier's precision/recall andfalse-green rate (node verifier-eval/run-eval.mjs ). This is how "is the verifier trustworthy?" becomes a number instead of a hope.

  • repo-verify/ — the product surface: point the verifier at awhole real repo . Archetype detection + the greenfieldstart on-ramp, the verifier CLI, the calibration store (the pattern/instance wall, made physical), and the git-awarePR gate + drop-in GitHub Action (action.yml ). Start atrepo-verify/README.md .

Business Source License 1.1 — see LICENSE. Free to use and self-host for any purpose, including commercially; the one reserved right is offering Forespec itself as a competing hosted service. Each released version converts to Apache 2.0 four years after it ships. The commitment it encodes: the local core is free and fully useful forever — never crippled to force an upgrade, no dark patterns — while paid, hosted plans fund the ongoing work that keeps the standard trustworthy; priced fairly, no lock-in.

── more in #ai-tools 4 stories · sorted by recency
── more on @forespec 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/forespec-catches-wha…] indexed:0 read:12min 2026-09-16 ·