cd /news/developer-tools/intuition-probe-let-an-llm-hallucina… · home topics developer-tools article
[ARTICLE · art-62020] src=gist.github.com ↗ pub= topic=developer-tools verified=true sentiment=· neutral

intuition-probe: let an LLM hallucinate the API/UX it expects, then conform your design to the guess. Companion to 'Complicated Thoughts On LLMs' (theocharis.dev)

A developer introduced intuition-probe, a tool that uses an LLM to guess the expected shape of an API, config, CLI, or UI before revealing the real design, then recommends conforming the design to the guess. The tool follows the principle 'make the API whatever the LLM guesses' to surface familiar patterns and improve adoption.

read4 min views35 publishedJul 14, 2026
name intuition-probe
description Use when you want to know what shape a developer or user reaches for when they first meet your API, config format, CLI, or UI, so you can make your design that shape. Spawns a blind agent that commits the design it EXPECTS before seeing the real thing, then reports the shape they reached for and conform-first recommendations (make the API whatever the guess is). Triggers on "is this intuitive", "would a developer guess this", "intuition probe", "blind-test this API/config/CLI", "test the DX/UX of", "what would someone expect here", "make the API whatever the LLM guesses".

Discover the shape a developer (or an LLM) reaches for when they meet your API, config, CLI, or UI for the first time, then make your design that shape. An agent guesses how to use the artifact blind, before it sees the real thing. The guess is not an error to grade; it is a design proposal. Where the guess diverges from your API, the default move is to conform the API to the guess, because the guess is what feels familiar, and familiar is what gets adopted.

This follows the Jazz principle "make the API whatever the LLM guesses": you only hear from the few users who push through an unfamiliar design, but an LLM shows you what people reach for cheaply and at volume. See references/going-deeper.md

for the theory, the conform-first default, and when NOT to conform.

It surfaces what the model's prior reaches for, the best cheap proxy for what is familiar to a developer. At the default N=1 it surfaces CANDIDATES, it does not CONFIRM: a confident divergent guess is a candidate shape to conform to, but a single draw cannot prove it is the familiar default, and a single match does not prove your API is familiar. Dial N up; only convergence across agents confirms a shape worth conforming to.

Step order matters: freeze prompt (3) before reading the artifact (4) before spawning agents (5). Reversing 3 and 4 contaminates the blind test.

Ask the user for: the artifact under test (a file path, package name, CLI command, or URL -- e.g. plugins/foo/SKILL.md

or git commit

), a rough goal, an optional read-set (a README, a docs URL, or a snippet), and N (default 1). If the artifact is a path/package/command, note it; do not read it yet.

Restate the rough goal as a pure outcome. Strip any leaked identifiers: method names, config keys, flags, exact labels. Show the user:

Sanitized goal: "" is that the task? (y / edit) Wait for confirmation. The read-set is NEVER sanitized (a README is allowed to contain the answer; that's the doc-informed test).

Take the template in references/blind-agent-prompt.md

and fill its placeholders (sanitized goal, system name, read-set if any) into a frozen working prompt that you hold in your own context. Do NOT edit the shipped template file. Also record N and the mode (cold / doc-informed) next to the frozen prompt so step 8 has them. Freeze the prompt now and do not revise it after step 4, so the real implementation cannot leak back into it. If the read-set is a URL, you (the orchestrator) fetch it now and inline its content into the frozen prompt; the blind agent is forbidden from opening URLs, so it must receive the doc text, never a bare link.

Read the real surface from the repo: source/signatures (API), schema/examples

(config), `--help`

/command defs (CLI), component/route source (UI). If you cannot

read it, ask the user to paste the real surface. Keep it private; it never goes into the blind prompt.

Spawn N agents via the Agent tool, each with the frozen prompt text from step 3 as its prompt. Use

subagent_type: general-purpose . Collect each agent's decision_points

JSON.

(For N=1, one call. For N>1, issue the Agent calls in parallel.)

For each agent's guess, apply the `references/scorer-rubric.md`

classification against the answer key, either inline (you, the orchestrator) or by spawning a scorer Agent. Produce the findings

JSON. Each finding names the familiar anchor the agent reached for and a conform-first recommendation; the adoption signal is read from the agent's recorded confidence (and, at N>1, convergence). Do not invent signal strength from hindsight. If you spawn a separate scorer Agent rather than scoring inline, pass it three things: the agent's decision_points

JSON, the answer key from step 4, and references/scorer-rubric.md

.

Group findings by decision point across agents. Mark a finding priority only if at least 2 independent agents hit it. Single-agent findings are listed, de-prioritized.

Render per references/report-format.md

: lead with the shape the prior reached for (the candidate spec) and a conform-first recommendation per divergence. If N is 1, include the candidate banner and never report the design as a confirmed familiar default. Offer to save the report to a scratch/artifacts folder. Never write to the repo under test.

── more in #developer-tools 4 stories · sorted by recency
── more on @intuition-probe 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/intuition-probe-let-…] indexed:0 read:4min 2026-07-14 ·