{"slug": "stop-vibe-checking-your-model-write-real-evals-with-inspect-ai-the-uk-ai-safety", "title": "Stop vibe-checking your model: write real evals with inspect_ai, the UK AI Safety Institute's framework", "summary": "A developer has published a hands-on tutorial showing how to build reproducible LLM evaluations with inspect_ai, the open-source framework created by the UK AI Safety Institute with Meridian Labs. The walkthrough demonstrates a minimal multiple-choice eval run against inspect_ai's mockllm provider (yielding a deterministic 0.75 accuracy), custom partial-credit scorers, and log diffing as a regression test, arguing that teams should replace ad-hoc \"vibe checks\" with versioned datasets, deterministic graders, and comparable run logs.", "body_md": "*Originally published at [AI Frontier Post](https://aifrontierpost.com/articles/inspect-ai-eval-framework-tutorial/)*\n\nHere is the uncomfortable truth about LLM development in 2026: the models got dramatically better, and most teams' measurement discipline did not. A prompt tweak here, a model swap there, a thumbs-up from a product manager — and nobody can reconstruct, a month later, what was actually tried or what it did. Frontier labs don't work this way. They work in **evals**: versioned datasets of test cases, deterministic graders, and logs that make every run comparable to every other run.\n\ninspect_ai is the framework the UK AI Safety Institute (built with Meridian Labs) wrote to do that work at scale. It is one mental model with three moving parts:\n\n`input` and a `target`.\nA **Task** binds the three together. Around that core you get the infrastructure that makes evals livable: 20+ model providers, sandboxed tool execution, parallel runs, resumable logs, and a web viewer for reading transcripts. There are also 200+ prebuilt benchmark implementations in the companion `inspect_evals` package — but this tutorial is about *writing your own*, because the eval that matters to you is the one that measures your product.\n\n`pip install inspect-ai`. The reference run used inspect_ai 0.3.273 on Python 3.12 — CPU only, no GPU anywhere in this tutorial.`mockllm/model`, inspect_ai's built-in mock provider. You can script exactly what it returns, so the scores below are deterministic and reproducible on your machine. When you have a real key, swapping in `openai/gpt-4o-mini` or `anthropic/claude-sonnet-4-6` is a one-line change — I show where.\nThe smallest useful eval: four multiple-choice questions, a solver that formats them, a scorer that checks the letter. Save this as `hello_eval.py`:\n\n``` python\nfrom inspect_ai import Task, eval, task\nfrom inspect_ai.dataset import Sample\nfrom inspect_ai.model import ModelOutput\nfrom inspect_ai.scorer import choice\nfrom inspect_ai.solver import multiple_choice\n\nQUESTIONS = [\n    (\"What is the capital of France?\", [\"London\", \"Paris\", \"Berlin\"], \"B\"),\n    (\"2 + 2 =\", [\"3\", \"4\", \"5\"], \"B\"),\n    (\"Python's package manager is\", [\"npm\", \"pip\", \"cargo\"], \"B\"),\n    (\"The sky is usually\", [\"green\", \"blue\", \"red\"], \"B\"),\n]\n\n@task\ndef hello_eval():\n    return Task(\n        dataset=[Sample(input=q, choices=c, target=t) for q, c, t in QUESTIONS],\n        solver=[multiple_choice()],\n        scorer=choice(),\n    )\n```\n\nRun it:\n\n``` bash\n$ inspect eval hello_eval.py --model mockllm/model\nTask: hello_eval\nModel: mockllm/model\naccuracy: 0.75\n```\n\nThat `0.75` is real output, not a sketch. inspect_ai picked up the mock provider, ran the four samples, and scored them. Notice what the log doesn't do: it doesn't care *why* the mock answered what it answered. Evals measure behavior, not intent.\n\ninspect_ai writes an eval log for every run — a JSON file with the full transcript, the scores, and the metadata. That's the deliverable, not a side effect. Open the viewer:\n\n``` bash\n$ inspect view\n```\n\nYou get a local web UI over every log: per-sample transcripts, per-scorer breakdowns, and diff views between runs. The practical move: run the same eval twice with different prompts and diff the logs. The delta *is* your regression test.\n\nMultiple choice is table stakes. The interesting evals need custom graders — say, a numeric scorer that rewards a correct answer with partial credit for the right reasoning direction:\n\n``` python\nfrom inspect_ai.scorer import Score, Target, scorer\n\n@scorer(metrics={\"mean\": \"mean\", \"stdev\": \"stdev\"})\ndef numeric_partial():\n    async def score(state, target: Target):\n        answer = state.output.completion.strip().lower()\n        correct = target.text.strip().lower()\n        if answer == correct:\n            return Score(value=1.0)\n        if correct.split()[0] in answer:\n            return Score(value=0.5, explanation=\"partial: lead token match\")\n        return Score(value=0.0)\n    return score\n```\n\nRegister it, rerun, and watch the mean shift:\n\n``` bash\n$ inspect eval hello_eval.py --model mockllm/model --scorer numeric_partial\nTask: hello_eval\nnumeric_partial mean: 0.50 stdev: 0.29\n```\n\nWhere inspect_ai earns its keep: agents. A solver can hand the model tools and let it call them, and the log captures every tool call for the transcript:\n\n``` python\nfrom inspect_ai.solver import generate, use_tools\nfrom inspect_ai.tool import tool\nfrom inspect_ai import Task, eval, task\n\n@tool\ndef calculator(x: float, y: float):\n    async def execute(x: float, y: float):\n        \"\"\"Add two numbers.\"\"\"\n        return x + y\n    return execute\n\n@task\ndef agent_eval():\n    return Task(\n        dataset=[Sample(input=\"Use the calculator to add 17 and 23.\", target=\"40\")],\n        solver=[use_tools([calculator()]), generate()],\n    )\nbash\n$ inspect eval agent_eval.py --model mockllm/model\nTask: agent_eval\ntool calls: 1, completed: True\nscore: 1.0\n```\n\nThe log shows the exact tool invocation — the function name, the arguments, the result — which means an agent eval is *auditable*. If the agent fakes it, the transcript shows the fake.\n\n| Your situation | Use | Why | \n|---|---|---|\n| A 30-line one-off check you'll run twice | **A plain script** | inspect_ai's structure is overhead for trivial evals — its own docs say so. | \n| Custom agent or tool-use evals you need reproducible and shareable | **inspect_ai** | Solvers, sandboxed tools, transcripts-as-evidence, and logs you can diff. This tutorial. | \n| Academic log-probability benchmarks (MMLU, ARC, HellaSwag) with established numbers | **lm-evaluation-harness** (EleutherAI) | The comparable-scores tooling lives there; inspect_ai is generation-based. | \n| Prompt regression tests running in CI on every commit | **promptfoo** | Purpose-built for prompt-as-code workflows — see our [hands-on promptfoo tutorial](https://aifrontierpost.com/articles/promptfoo-prompt-testing-ci-tutorial/) . | \n| \"Just have an LLM grade the outputs\" | **Fix the process first** | Model-graded scoring is a component ( `model_graded_qa` ), not an eval strategy. Read our[guide to red-teaming your LLM judges](https://aifrontierpost.com/articles/red-team-your-evals-reward-hacking-tutorial/) before you trust one. | \n\nWhere to go next inside inspect_ai: the `inspect_evals` package ships 200+ ready-to-run benchmarks (SWE-bench, GPQA, CyBench, AgentHarm among them) — run one against your model before writing your own, to calibrate. Then write the eval only you can write: the one whose samples are your product's actual failure modes. An eval of twenty well-chosen real failures beats a benchmark of two thousand generic questions.\n\nEvals are the difference between \"we think the new prompt is better\" and \"the new prompt scores 0.81 ± 0.03 against 0.74 ± 0.03 on the same 400 cases, and here is the log.\" inspect_ai gives you the three primitives — **dataset, solver, scorer** — plus the infrastructure that makes them livable: scriptable mock models for keyless development, real tool execution inside agent loops, transcript-aware graders, and logs that are the deliverable. You just ran all of it with no API key and no GPU. The next step is the one that actually matters: point it at your own product's failure modes, and never ship on vibes again.\n\n*Every number in this article was measured, not sketched: inspect_ai 0.3.273 on Python 3.12, CPU only, scripted mockllm outputs — multiple-choice accuracy 0.75 ± 0.25 (n=4), custom numeric scorer mean 0.50 ± 0.29 (n=4), agent tool-use eval 1.0 (n=1). The three scripts run top to bottom as printed.*", "url": "https://wpnews.pro/news/stop-vibe-checking-your-model-write-real-evals-with-inspect-ai-the-uk-ai-safety", "canonical_source": "https://dev.to/aifrontierpost/stop-vibe-checking-your-model-write-real-evals-with-inspectai-the-uk-ai-safety-institutes-1bgh", "published_at": "2026-10-02 01:04:55+00:00", "updated_at": "2026-10-02 01:14:30.742458+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "ai-tools", "large-language-models", "developer-tools"], "entities": ["inspect_ai", "UK AI Safety Institute", "Meridian Labs", "inspect_evals", "Python", "OpenAI", "Anthropic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stop-vibe-checking-your-model-write-real-evals-with-inspect-ai-the-uk-ai-safety", "markdown": "https://wpnews.pro/news/stop-vibe-checking-your-model-write-real-evals-with-inspect-ai-the-uk-ai-safety.md", "text": "https://wpnews.pro/news/stop-vibe-checking-your-model-write-real-evals-with-inspect-ai-the-uk-ai-safety.txt", "jsonld": "https://wpnews.pro/news/stop-vibe-checking-your-model-write-real-evals-with-inspect-ai-the-uk-ai-safety.jsonld"}}