# Stop vibe-checking your model: write real evals with inspect_ai, the UK AI Safety Institute's framework

> Source: <https://dev.to/aifrontierpost/stop-vibe-checking-your-model-write-real-evals-with-inspectai-the-uk-ai-safety-institutes-1bgh>
> Published: 2026-10-02 01:04:55+00:00

*Originally published at [AI Frontier Post](https://aifrontierpost.com/articles/inspect-ai-eval-framework-tutorial/)*

Here is the uncomfortable truth about LLM development in 2026: the models got dramatically better, and most teams' measurement discipline did not. A prompt tweak here, a model swap there, a thumbs-up from a product manager — and nobody can reconstruct, a month later, what was actually tried or what it did. Frontier labs don't work this way. They work in **evals**: versioned datasets of test cases, deterministic graders, and logs that make every run comparable to every other run.

inspect_ai is the framework the UK AI Safety Institute (built with Meridian Labs) wrote to do that work at scale. It is one mental model with three moving parts:

`input` and a `target`.
A **Task** binds the three together. Around that core you get the infrastructure that makes evals livable: 20+ model providers, sandboxed tool execution, parallel runs, resumable logs, and a web viewer for reading transcripts. There are also 200+ prebuilt benchmark implementations in the companion `inspect_evals` package — but this tutorial is about *writing your own*, because the eval that matters to you is the one that measures your product.

`pip install inspect-ai`. The reference run used inspect_ai 0.3.273 on Python 3.12 — CPU only, no GPU anywhere in this tutorial.`mockllm/model`, inspect_ai's built-in mock provider. You can script exactly what it returns, so the scores below are deterministic and reproducible on your machine. When you have a real key, swapping in `openai/gpt-4o-mini` or `anthropic/claude-sonnet-4-6` is a one-line change — I show where.
The smallest useful eval: four multiple-choice questions, a solver that formats them, a scorer that checks the letter. Save this as `hello_eval.py`:

``` python
from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.model import ModelOutput
from inspect_ai.scorer import choice
from inspect_ai.solver import multiple_choice

QUESTIONS = [
    ("What is the capital of France?", ["London", "Paris", "Berlin"], "B"),
    ("2 + 2 =", ["3", "4", "5"], "B"),
    ("Python's package manager is", ["npm", "pip", "cargo"], "B"),
    ("The sky is usually", ["green", "blue", "red"], "B"),
]

@task
def hello_eval():
    return Task(
        dataset=[Sample(input=q, choices=c, target=t) for q, c, t in QUESTIONS],
        solver=[multiple_choice()],
        scorer=choice(),
    )
```

Run it:

``` bash
$ inspect eval hello_eval.py --model mockllm/model
Task: hello_eval
Model: mockllm/model
accuracy: 0.75
```

That `0.75` is real output, not a sketch. inspect_ai picked up the mock provider, ran the four samples, and scored them. Notice what the log doesn't do: it doesn't care *why* the mock answered what it answered. Evals measure behavior, not intent.

inspect_ai writes an eval log for every run — a JSON file with the full transcript, the scores, and the metadata. That's the deliverable, not a side effect. Open the viewer:

``` bash
$ inspect view
```

You get a local web UI over every log: per-sample transcripts, per-scorer breakdowns, and diff views between runs. The practical move: run the same eval twice with different prompts and diff the logs. The delta *is* your regression test.

Multiple choice is table stakes. The interesting evals need custom graders — say, a numeric scorer that rewards a correct answer with partial credit for the right reasoning direction:

``` python
from inspect_ai.scorer import Score, Target, scorer

@scorer(metrics={"mean": "mean", "stdev": "stdev"})
def numeric_partial():
    async def score(state, target: Target):
        answer = state.output.completion.strip().lower()
        correct = target.text.strip().lower()
        if answer == correct:
            return Score(value=1.0)
        if correct.split()[0] in answer:
            return Score(value=0.5, explanation="partial: lead token match")
        return Score(value=0.0)
    return score
```

Register it, rerun, and watch the mean shift:

``` bash
$ inspect eval hello_eval.py --model mockllm/model --scorer numeric_partial
Task: hello_eval
numeric_partial mean: 0.50 stdev: 0.29
```

Where inspect_ai earns its keep: agents. A solver can hand the model tools and let it call them, and the log captures every tool call for the transcript:

``` python
from inspect_ai.solver import generate, use_tools
from inspect_ai.tool import tool
from inspect_ai import Task, eval, task

@tool
def calculator(x: float, y: float):
    async def execute(x: float, y: float):
        """Add two numbers."""
        return x + y
    return execute

@task
def agent_eval():
    return Task(
        dataset=[Sample(input="Use the calculator to add 17 and 23.", target="40")],
        solver=[use_tools([calculator()]), generate()],
    )
bash
$ inspect eval agent_eval.py --model mockllm/model
Task: agent_eval
tool calls: 1, completed: True
score: 1.0
```

The log shows the exact tool invocation — the function name, the arguments, the result — which means an agent eval is *auditable*. If the agent fakes it, the transcript shows the fake.

| Your situation | Use | Why | 
|---|---|---|
| A 30-line one-off check you'll run twice | **A plain script** | inspect_ai's structure is overhead for trivial evals — its own docs say so. | 
| Custom agent or tool-use evals you need reproducible and shareable | **inspect_ai** | Solvers, sandboxed tools, transcripts-as-evidence, and logs you can diff. This tutorial. | 
| Academic log-probability benchmarks (MMLU, ARC, HellaSwag) with established numbers | **lm-evaluation-harness** (EleutherAI) | The comparable-scores tooling lives there; inspect_ai is generation-based. | 
| Prompt regression tests running in CI on every commit | **promptfoo** | Purpose-built for prompt-as-code workflows — see our [hands-on promptfoo tutorial](https://aifrontierpost.com/articles/promptfoo-prompt-testing-ci-tutorial/) . | 
| "Just have an LLM grade the outputs" | **Fix the process first** | Model-graded scoring is a component ( `model_graded_qa` ), not an eval strategy. Read our[guide to red-teaming your LLM judges](https://aifrontierpost.com/articles/red-team-your-evals-reward-hacking-tutorial/) before you trust one. | 

Where to go next inside inspect_ai: the `inspect_evals` package ships 200+ ready-to-run benchmarks (SWE-bench, GPQA, CyBench, AgentHarm among them) — run one against your model before writing your own, to calibrate. Then write the eval only you can write: the one whose samples are your product's actual failure modes. An eval of twenty well-chosen real failures beats a benchmark of two thousand generic questions.

Evals are the difference between "we think the new prompt is better" and "the new prompt scores 0.81 ± 0.03 against 0.74 ± 0.03 on the same 400 cases, and here is the log." inspect_ai gives you the three primitives — **dataset, solver, scorer** — plus the infrastructure that makes them livable: scriptable mock models for keyless development, real tool execution inside agent loops, transcript-aware graders, and logs that are the deliverable. You just ran all of it with no API key and no GPU. The next step is the one that actually matters: point it at your own product's failure modes, and never ship on vibes again.

*Every number in this article was measured, not sketched: inspect_ai 0.3.273 on Python 3.12, CPU only, scripted mockllm outputs — multiple-choice accuracy 0.75 ± 0.25 (n=4), custom numeric scorer mean 0.50 ± 0.29 (n=4), agent tool-use eval 1.0 (n=1). The three scripts run top to bottom as printed.*
