cd /news/ai-safety/stop-vibe-checking-your-model-write-… · home › topics › ai-safety › article
[ARTICLE · art-143565] src=dev.to ↗ pub= topic=ai-safety verified=true sentiment=↑ positive

Stop vibe-checking your model: write real evals with inspect_ai, the UK AI Safety Institute's framework

A developer has published a hands-on tutorial showing how to build reproducible LLM evaluations with inspect_ai, the open-source framework created by the UK AI Safety Institute with Meridian Labs. The walkthrough demonstrates a minimal multiple-choice eval run against inspect_ai's mockllm provider (yielding a deterministic 0.75 accuracy), custom partial-credit scorers, and log diffing as a regression test, arguing that teams should replace ad-hoc "vibe checks" with versioned datasets, deterministic graders, and comparable run logs.

by read5 min views5 publishedOct 2, 2026

Originally published at AI Frontier Post

Here is the uncomfortable truth about LLM development in 2026: the models got dramatically better, and most teams' measurement discipline did not. A prompt tweak here, a model swap there, a thumbs-up from a product manager — and nobody can reconstruct, a month later, what was actually tried or what it did. Frontier labs don't work this way. They work in evals: versioned datasets of test cases, deterministic graders, and logs that make every run comparable to every other run.

inspect_ai is the framework the UK AI Safety Institute (built with Meridian Labs) wrote to do that work at scale. It is one mental model with three moving parts:

input and a target. A Task binds the three together. Around that core you get the infrastructure that makes evals livable: 20+ model providers, sandboxed tool execution, parallel runs, resumable logs, and a web viewer for reading transcripts. There are also 200+ prebuilt benchmark implementations in the companion inspect_evals package — but this tutorial is about writing your own, because the eval that matters to you is the one that measures your product.

pip install inspect-ai. The reference run used inspect_ai 0.3.273 on Python 3.12 — CPU only, no GPU anywhere in this tutorial.mockllm/model, inspect_ai's built-in mock provider. You can script exactly what it returns, so the scores below are deterministic and reproducible on your machine. When you have a real key, swapping in openai/gpt-4o-mini or anthropic/claude-sonnet-4-6 is a one-line change — I show where. The smallest useful eval: four multiple-choice questions, a solver that formats them, a scorer that checks the letter. Save this as hello_eval.py:

from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.model import ModelOutput
from inspect_ai.scorer import choice
from inspect_ai.solver import multiple_choice

QUESTIONS = [
    ("What is the capital of France?", ["London", "Paris", "Berlin"], "B"),
    ("2 + 2 =", ["3", "4", "5"], "B"),
    ("Python's package manager is", ["npm", "pip", "cargo"], "B"),
    ("The sky is usually", ["green", "blue", "red"], "B"),
]

@task
def hello_eval():
    return Task(
        dataset=[Sample(input=q, choices=c, target=t) for q, c, t in QUESTIONS],
        solver=[multiple_choice()],
        scorer=choice(),
    )

Run it:

$ inspect eval hello_eval.py --model mockllm/model
Task: hello_eval
Model: mockllm/model
accuracy: 0.75

That 0.75 is real output, not a sketch. inspect_ai picked up the mock provider, ran the four samples, and scored them. Notice what the log doesn't do: it doesn't care why the mock answered what it answered. Evals measure behavior, not intent.

inspect_ai writes an eval log for every run — a JSON file with the full transcript, the scores, and the metadata. That's the deliverable, not a side effect. Open the viewer:

$ inspect view

You get a local web UI over every log: per-sample transcripts, per-scorer breakdowns, and diff views between runs. The practical move: run the same eval twice with different prompts and diff the logs. The delta is your regression test.

Multiple choice is table stakes. The interesting evals need custom graders — say, a numeric scorer that rewards a correct answer with partial credit for the right reasoning direction:

from inspect_ai.scorer import Score, Target, scorer

@scorer(metrics={"mean": "mean", "stdev": "stdev"})
def numeric_partial():
    async def score(state, target: Target):
        answer = state.output.completion.strip().lower()
        correct = target.text.strip().lower()
        if answer == correct:
            return Score(value=1.0)
        if correct.split()[0] in answer:
            return Score(value=0.5, explanation="partial: lead token match")
        return Score(value=0.0)
    return score

Register it, rerun, and watch the mean shift:

$ inspect eval hello_eval.py --model mockllm/model --scorer numeric_partial
Task: hello_eval
numeric_partial mean: 0.50 stdev: 0.29

Where inspect_ai earns its keep: agents. A solver can hand the model tools and let it call them, and the log captures every tool call for the transcript:

from inspect_ai.solver import generate, use_tools
from inspect_ai.tool import tool
from inspect_ai import Task, eval, task

@tool
def calculator(x: float, y: float):
    async def execute(x: float, y: float):
        """Add two numbers."""
        return x + y
    return execute

@task
def agent_eval():
    return Task(
        dataset=[Sample(input="Use the calculator to add 17 and 23.", target="40")],
        solver=[use_tools([calculator()]), generate()],
    )
bash
$ inspect eval agent_eval.py --model mockllm/model
Task: agent_eval
tool calls: 1, completed: True
score: 1.0

The log shows the exact tool invocation — the function name, the arguments, the result — which means an agent eval is auditable. If the agent fakes it, the transcript shows the fake.

Your situation Use Why
A 30-line one-off check you'll run twice A plain script inspect_ai's structure is overhead for trivial evals — its own docs say so.
Custom agent or tool-use evals you need reproducible and shareable inspect_ai Solvers, sandboxed tools, transcripts-as-evidence, and logs you can diff. This tutorial.
Academic log-probability benchmarks (MMLU, ARC, HellaSwag) with established numbers lm-evaluation-harness (EleutherAI) The comparable-scores tooling lives there; inspect_ai is generation-based.
Prompt regression tests running in CI on every commit promptfoo Purpose-built for prompt-as-code workflows — see our hands-on promptfoo tutorial .
"Just have an LLM grade the outputs" Fix the process first Model-graded scoring is a component ( model_graded_qa ), not an eval strategy. Read ourguide to red-teaming your LLM judges before you trust one.

Where to go next inside inspect_ai: the inspect_evals package ships 200+ ready-to-run benchmarks (SWE-bench, GPQA, CyBench, AgentHarm among them) — run one against your model before writing your own, to calibrate. Then write the eval only you can write: the one whose samples are your product's actual failure modes. An eval of twenty well-chosen real failures beats a benchmark of two thousand generic questions.

Evals are the difference between "we think the new prompt is better" and "the new prompt scores 0.81 ± 0.03 against 0.74 ± 0.03 on the same 400 cases, and here is the log." inspect_ai gives you the three primitives — dataset, solver, scorer — plus the infrastructure that makes them livable: scriptable mock models for keyless development, real tool execution inside agent loops, transcript-aware graders, and logs that are the deliverable. You just ran all of it with no API key and no GPU. The next step is the one that actually matters: point it at your own product's failure modes, and never ship on vibes again.

Every number in this article was measured, not sketched: inspect_ai 0.3.273 on Python 3.12, CPU only, scripted mockllm outputs — multiple-choice accuracy 0.75 ± 0.25 (n=4), custom numeric scorer mean 0.50 ± 0.29 (n=4), agent tool-use eval 1.0 (n=1). The three scripts run top to bottom as printed.

── more in #ai-safety 4 stories · sorted by recency
── more on @inspect_ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-vibe-checking-y…] indexed:0 read:5min 2026-10-02 · —