Stop vibe-checking your model: write real evals with inspect_ai, the UK AI Safety Institute's framework A developer has published a hands-on tutorial showing how to build reproducible LLM evaluations with inspect_ai, the open-source framework created by the UK AI Safety Institute with Meridian Labs. The walkthrough demonstrates a minimal multiple-choice eval run against inspect_ai's mockllm provider (yielding a deterministic 0.75 accuracy), custom partial-credit scorers, and log diffing as a regression test, arguing that teams should replace ad-hoc "vibe checks" with versioned datasets, deterministic graders, and comparable run logs. Originally published at AI Frontier Post https://aifrontierpost.com/articles/inspect-ai-eval-framework-tutorial/ Here is the uncomfortable truth about LLM development in 2026: the models got dramatically better, and most teams' measurement discipline did not. A prompt tweak here, a model swap there, a thumbs-up from a product manager — and nobody can reconstruct, a month later, what was actually tried or what it did. Frontier labs don't work this way. They work in evals : versioned datasets of test cases, deterministic graders, and logs that make every run comparable to every other run. inspect ai is the framework the UK AI Safety Institute built with Meridian Labs wrote to do that work at scale. It is one mental model with three moving parts: input and a target . A Task binds the three together. Around that core you get the infrastructure that makes evals livable: 20+ model providers, sandboxed tool execution, parallel runs, resumable logs, and a web viewer for reading transcripts. There are also 200+ prebuilt benchmark implementations in the companion inspect evals package — but this tutorial is about writing your own , because the eval that matters to you is the one that measures your product. pip install inspect-ai . The reference run used inspect ai 0.3.273 on Python 3.12 — CPU only, no GPU anywhere in this tutorial. mockllm/model , inspect ai's built-in mock provider. You can script exactly what it returns, so the scores below are deterministic and reproducible on your machine. When you have a real key, swapping in openai/gpt-4o-mini or anthropic/claude-sonnet-4-6 is a one-line change — I show where. The smallest useful eval: four multiple-choice questions, a solver that formats them, a scorer that checks the letter. Save this as hello eval.py : python from inspect ai import Task, eval, task from inspect ai.dataset import Sample from inspect ai.model import ModelOutput from inspect ai.scorer import choice from inspect ai.solver import multiple choice QUESTIONS = "What is the capital of France?", "London", "Paris", "Berlin" , "B" , "2 + 2 =", "3", "4", "5" , "B" , "Python's package manager is", "npm", "pip", "cargo" , "B" , "The sky is usually", "green", "blue", "red" , "B" , @task def hello eval : return Task dataset= Sample input=q, choices=c, target=t for q, c, t in QUESTIONS , solver= multiple choice , scorer=choice , Run it: bash $ inspect eval hello eval.py --model mockllm/model Task: hello eval Model: mockllm/model accuracy: 0.75 That 0.75 is real output, not a sketch. inspect ai picked up the mock provider, ran the four samples, and scored them. Notice what the log doesn't do: it doesn't care why the mock answered what it answered. Evals measure behavior, not intent. inspect ai writes an eval log for every run — a JSON file with the full transcript, the scores, and the metadata. That's the deliverable, not a side effect. Open the viewer: bash $ inspect view You get a local web UI over every log: per-sample transcripts, per-scorer breakdowns, and diff views between runs. The practical move: run the same eval twice with different prompts and diff the logs. The delta is your regression test. Multiple choice is table stakes. The interesting evals need custom graders — say, a numeric scorer that rewards a correct answer with partial credit for the right reasoning direction: python from inspect ai.scorer import Score, Target, scorer @scorer metrics={"mean": "mean", "stdev": "stdev"} def numeric partial : async def score state, target: Target : answer = state.output.completion.strip .lower correct = target.text.strip .lower if answer == correct: return Score value=1.0 if correct.split 0 in answer: return Score value=0.5, explanation="partial: lead token match" return Score value=0.0 return score Register it, rerun, and watch the mean shift: bash $ inspect eval hello eval.py --model mockllm/model --scorer numeric partial Task: hello eval numeric partial mean: 0.50 stdev: 0.29 Where inspect ai earns its keep: agents. A solver can hand the model tools and let it call them, and the log captures every tool call for the transcript: python from inspect ai.solver import generate, use tools from inspect ai.tool import tool from inspect ai import Task, eval, task @tool def calculator x: float, y: float : async def execute x: float, y: float : """Add two numbers.""" return x + y return execute @task def agent eval : return Task dataset= Sample input="Use the calculator to add 17 and 23.", target="40" , solver= use tools calculator , generate , bash $ inspect eval agent eval.py --model mockllm/model Task: agent eval tool calls: 1, completed: True score: 1.0 The log shows the exact tool invocation — the function name, the arguments, the result — which means an agent eval is auditable . If the agent fakes it, the transcript shows the fake. | Your situation | Use | Why | |---|---|---| | A 30-line one-off check you'll run twice | A plain script | inspect ai's structure is overhead for trivial evals — its own docs say so. | | Custom agent or tool-use evals you need reproducible and shareable | inspect ai | Solvers, sandboxed tools, transcripts-as-evidence, and logs you can diff. This tutorial. | | Academic log-probability benchmarks MMLU, ARC, HellaSwag with established numbers | lm-evaluation-harness EleutherAI | The comparable-scores tooling lives there; inspect ai is generation-based. | | Prompt regression tests running in CI on every commit | promptfoo | Purpose-built for prompt-as-code workflows — see our hands-on promptfoo tutorial https://aifrontierpost.com/articles/promptfoo-prompt-testing-ci-tutorial/ . | | "Just have an LLM grade the outputs" | Fix the process first | Model-graded scoring is a component model graded qa , not an eval strategy. Read our guide to red-teaming your LLM judges https://aifrontierpost.com/articles/red-team-your-evals-reward-hacking-tutorial/ before you trust one. | Where to go next inside inspect ai: the inspect evals package ships 200+ ready-to-run benchmarks SWE-bench, GPQA, CyBench, AgentHarm among them — run one against your model before writing your own, to calibrate. Then write the eval only you can write: the one whose samples are your product's actual failure modes. An eval of twenty well-chosen real failures beats a benchmark of two thousand generic questions. Evals are the difference between "we think the new prompt is better" and "the new prompt scores 0.81 ± 0.03 against 0.74 ± 0.03 on the same 400 cases, and here is the log." inspect ai gives you the three primitives — dataset, solver, scorer — plus the infrastructure that makes them livable: scriptable mock models for keyless development, real tool execution inside agent loops, transcript-aware graders, and logs that are the deliverable. You just ran all of it with no API key and no GPU. The next step is the one that actually matters: point it at your own product's failure modes, and never ship on vibes again. Every number in this article was measured, not sketched: inspect ai 0.3.273 on Python 3.12, CPU only, scripted mockllm outputs — multiple-choice accuracy 0.75 ± 0.25 n=4 , custom numeric scorer mean 0.50 ± 0.29 n=4 , agent tool-use eval 1.0 n=1 . The three scripts run top to bottom as printed.