Originally published at AI Frontier Post
Here is the uncomfortable truth about LLM development in 2026: the models got dramatically better, and most teams' measurement discipline did not. A prompt tweak here, a model swap there, a thumbs-up from a product manager — and nobody can reconstruct, a month later, what was actually tried or what it did. Frontier labs don't work this way. They work in evals: versioned datasets of test cases, deterministic graders, and logs that make every run comparable to every other run.
inspect_ai is the framework the UK AI Safety Institute (built with Meridian Labs) wrote to do that work at scale. It is one mental model with three moving parts:
input and a target.
A Task binds the three together. Around that core you get the infrastructure that makes evals livable: 20+ model providers, sandboxed tool execution, parallel runs, resumable logs, and a web viewer for reading transcripts. There are also 200+ prebuilt benchmark implementations in the companion inspect_evals package — but this tutorial is about writing your own, because the eval that matters to you is the one that measures your product.
pip install inspect-ai. The reference run used inspect_ai 0.3.273 on Python 3.12 — CPU only, no GPU anywhere in this tutorial.mockllm/model, inspect_ai's built-in mock provider. You can script exactly what it returns, so the scores below are deterministic and reproducible on your machine. When you have a real key, swapping in openai/gpt-4o-mini or anthropic/claude-sonnet-4-6 is a one-line change — I show where.
The smallest useful eval: four multiple-choice questions, a solver that formats them, a scorer that checks the letter. Save this as hello_eval.py:
from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.model import ModelOutput
from inspect_ai.scorer import choice
from inspect_ai.solver import multiple_choice
QUESTIONS = [
("What is the capital of France?", ["London", "Paris", "Berlin"], "B"),
("2 + 2 =", ["3", "4", "5"], "B"),
("Python's package manager is", ["npm", "pip", "cargo"], "B"),
("The sky is usually", ["green", "blue", "red"], "B"),
]
@task
def hello_eval():
return Task(
dataset=[Sample(input=q, choices=c, target=t) for q, c, t in QUESTIONS],
solver=[multiple_choice()],
scorer=choice(),
)
Run it:
$ inspect eval hello_eval.py --model mockllm/model
Task: hello_eval
Model: mockllm/model
accuracy: 0.75
That 0.75 is real output, not a sketch. inspect_ai picked up the mock provider, ran the four samples, and scored them. Notice what the log doesn't do: it doesn't care why the mock answered what it answered. Evals measure behavior, not intent.
inspect_ai writes an eval log for every run — a JSON file with the full transcript, the scores, and the metadata. That's the deliverable, not a side effect. Open the viewer:
$ inspect view
You get a local web UI over every log: per-sample transcripts, per-scorer breakdowns, and diff views between runs. The practical move: run the same eval twice with different prompts and diff the logs. The delta is your regression test.
Multiple choice is table stakes. The interesting evals need custom graders — say, a numeric scorer that rewards a correct answer with partial credit for the right reasoning direction:
from inspect_ai.scorer import Score, Target, scorer
@scorer(metrics={"mean": "mean", "stdev": "stdev"})
def numeric_partial():
async def score(state, target: Target):
answer = state.output.completion.strip().lower()
correct = target.text.strip().lower()
if answer == correct:
return Score(value=1.0)
if correct.split()[0] in answer:
return Score(value=0.5, explanation="partial: lead token match")
return Score(value=0.0)
return score
Register it, rerun, and watch the mean shift:
$ inspect eval hello_eval.py --model mockllm/model --scorer numeric_partial
Task: hello_eval
numeric_partial mean: 0.50 stdev: 0.29
Where inspect_ai earns its keep: agents. A solver can hand the model tools and let it call them, and the log captures every tool call for the transcript:
from inspect_ai.solver import generate, use_tools
from inspect_ai.tool import tool
from inspect_ai import Task, eval, task
@tool
def calculator(x: float, y: float):
async def execute(x: float, y: float):
"""Add two numbers."""
return x + y
return execute
@task
def agent_eval():
return Task(
dataset=[Sample(input="Use the calculator to add 17 and 23.", target="40")],
solver=[use_tools([calculator()]), generate()],
)
bash
$ inspect eval agent_eval.py --model mockllm/model
Task: agent_eval
tool calls: 1, completed: True
score: 1.0
The log shows the exact tool invocation — the function name, the arguments, the result — which means an agent eval is auditable. If the agent fakes it, the transcript shows the fake.
| Your situation | Use | Why |
|---|---|---|
| A 30-line one-off check you'll run twice | A plain script | inspect_ai's structure is overhead for trivial evals — its own docs say so. |
| Custom agent or tool-use evals you need reproducible and shareable | inspect_ai | Solvers, sandboxed tools, transcripts-as-evidence, and logs you can diff. This tutorial. |
| Academic log-probability benchmarks (MMLU, ARC, HellaSwag) with established numbers | lm-evaluation-harness (EleutherAI) | The comparable-scores tooling lives there; inspect_ai is generation-based. |
| Prompt regression tests running in CI on every commit | promptfoo | Purpose-built for prompt-as-code workflows — see our hands-on promptfoo tutorial . |
| "Just have an LLM grade the outputs" | Fix the process first | Model-graded scoring is a component ( model_graded_qa ), not an eval strategy. Read ourguide to red-teaming your LLM judges before you trust one. |
Where to go next inside inspect_ai: the inspect_evals package ships 200+ ready-to-run benchmarks (SWE-bench, GPQA, CyBench, AgentHarm among them) — run one against your model before writing your own, to calibrate. Then write the eval only you can write: the one whose samples are your product's actual failure modes. An eval of twenty well-chosen real failures beats a benchmark of two thousand generic questions.
Evals are the difference between "we think the new prompt is better" and "the new prompt scores 0.81 ± 0.03 against 0.74 ± 0.03 on the same 400 cases, and here is the log." inspect_ai gives you the three primitives — dataset, solver, scorer — plus the infrastructure that makes them livable: scriptable mock models for keyless development, real tool execution inside agent loops, transcript-aware graders, and logs that are the deliverable. You just ran all of it with no API key and no GPU. The next step is the one that actually matters: point it at your own product's failure modes, and never ship on vibes again.
Every number in this article was measured, not sketched: inspect_ai 0.3.273 on Python 3.12, CPU only, scripted mockllm outputs — multiple-choice accuracy 0.75 ± 0.25 (n=4), custom numeric scorer mean 0.50 ± 0.29 (n=4), agent tool-use eval 1.0 (n=1). The three scripts run top to bottom as printed.