Inspect AI in Production: A Hard-Nosed Review of Agent Evals, Logs, and Model-Upgrade Gates A developer evaluated Inspect AI as an execution and artifact layer for item-level agent-eval gates, concluding that it provides a stronger foundation than a home-grown request loop but leaves raw log reproducibility, scorer contracts, sandbox boundaries, and cost accounting as engineering responsibilities. The review, run with inspect-ai 0.3.100 on Python 3.12 in an isolated container, recommends pinning a small subset of the Inspect Evals corpus into an internal regression suite rather than importing a moving benchmark wholesale. The author notes that Inspect does not by itself make a benchmark deterministic, a model grader trustworthy, or a tool safe. The model upgrade itself is rarely the difficult part. The difficult part is proving that the new model did not silently damage extraction accuracy, tool behavior, refusal handling, or operational cost. Inspect AI Verdict Inspect AI provides the execution and artifact layer for item-level agent-eval gates, but raw log reproducibility, scorer contracts, sandbox boundaries, and cost accounting remain engineering responsibilities. We have seen the same internal harness emerge repeatedly: a JSONL dataset, an asynchronous request loop, a pile of retry logic, a scorer script, and a spreadsheet comparing the old model with the new one. It works until an incident forces us to answer questions the harness was never designed to answer: We evaluated Inspect AI https://inspect.aisi.org.uk/ because it handles that entire execution layer. An Inspect evaluation combines a dataset, a solver, and one or more scorers. The solver may be a single generation call, a custom function decorated with @solver , or a multi-turn agent. The scorer may perform exact matching, inclusion matching, model grading, or arbitrary custom logic. The framework also writes structured .eval logs and ships the inspect view browser interface. That separation matters. We do not want the dashboard to be the source of truth; we want the artifact to be the source of truth and the dashboard to be one way of reading it. The agent surface was another reason for testing it. Inspect accepts agents in the solver position and includes a ReAct loop, tool interfaces, message and token limits, custom agents, agent handoffs, and adapters for external agents. Its tool catalog includes shell, Python, text editing, web, computer, and MCP-oriented interfaces. For dangerous execution, it supports sandbox backends including Docker and more infrastructure-heavy options. We also inspected the Inspect Evals repository https://github.com/UKGovernmentBEIS/inspect evals rather than evaluating the framework only with toy arithmetic. It provides a substantial task corpus and demonstrates how real evaluations organize datasets, task versions, dependencies, scoring, and sandbox assets. We still recommend pinning a small subset into an internal regression suite rather than importing a moving benchmark wholesale. Our short conclusion before the details: Inspect gives us a much stronger foundation than another home-grown request loop. It does not, by itself, make a benchmark deterministic, a model grader trustworthy, or a tool safe. We ran the offline scorer check with Python 3.12 in an isolated container. For the local walkthrough, we recommend a virtual environment. That was deliberate: Python 3.11 and 3.12 are the Inspect Evals project's preferred targets, while later Python versions carry compatibility caveats for parts of the broader task corpus. For the local walkthrough, use this installation path: python3.12 -m venv .venv source .venv/bin/activate python -m pip install --upgrade pip python -m pip install "inspect-ai==0.3.100" openai export OPENAI API KEY="replace-me" mkdir -p logs The pinned inspect-ai wheel resolved successfully in our environment. The wheel we downloaded had this SHA-256 digest: 78e4c5b7e426fcc25563cd68f9f976a1c18c00adc9f4955bde9a6c97016363d7 For production, we would pin the complete lockfile rather than only the top-level package. Inspect pulls in a nontrivial dependency tree that includes asynchronous HTTP, schema, filesystem, terminal UI, and cloud-storage packages. We then reduced the task shape to the smallest useful production example: registered task, registered solver, custom partial-credit scorer, stable sample IDs, and an explicit log directory. python regression eval.py import os import re from inspect ai import Task, eval, task from inspect ai.dataset import Sample from inspect ai.scorer import Score, Target, scorer from inspect ai.solver import Generate, TaskState, solver @solver def answer once : async def solve state: TaskState, generate: Generate - TaskState: return await generate state return solve def normalize text: str - str: return re.sub r"\s+", " ", text.strip .lower @scorer def partial answer : async def score state: TaskState, target: Target - Score: answer = normalize state.output.completion expected = normalize target.text if answer == expected: value = 1.0 reason = "Exact normalized match" elif expected in answer: value = 0.5 reason = "Target present with additional text" else: value = 0.0 reason = "Target absent" return Score value=value, answer=state.output.completion, explanation=reason, return score @task def upgrade gate : return Task dataset= Sample id="refund-window", input="Reply with only the refund window: 30 days", target="30 days", , Sample id="support-tier", input="Reply with only the support tier: enterprise", target="enterprise", , , solver=answer once , scorer=partial answer , if name == " main ": eval upgrade gate , model=os.environ.get "INSPECT MODEL", "openai/gpt-5" , log dir="logs", To execute this worked example through Python or the CLI and inspect its logs, use: export INSPECT MODEL="openai/gpt-5" python regression eval.py Equivalent task discovery through the CLI inspect eval regression eval.py@upgrade gate \ --model "$INSPECT MODEL" \ --log-dir logs inspect view --log-dir logs A successful run should produce output in this general shape. The following console block is illustrative, not a measurement from our failed qualification run: bash $ inspect eval regression eval.py@upgrade gate \ --model openai/gpt-5 \ --log-dir logs upgrade gate 2 samples : complete Model: openai/gpt-5 Log: logs/2026-10-11T101530 upgrade-gate.eval Scores: partial answer: 0.750 Samples: completed: 2 errors: 0 For an agent task, the task definition changes more than the execution command. This is the minimal pattern we would use before adding a pinned challenge dataset: python from inspect ai import Task, task from inspect ai.agent import react from inspect ai.scorer import includes from inspect ai.tool import bash, python, todo write @task def sandbox agent eval : return Task dataset=load pinned samples , solver=react prompt="Use the available tools and submit only the final answer.", tools= bash , python , todo write , attempts=3, , scorer=includes , sandbox="docker", message limit=30, The important production property is not the decorator syntax. It is that sample inputs, targets, messages, tool calls, outputs, scores, and run configuration can travel together in one reviewable artifact. Our first narrow scorer contract test did not complete successfully. We created five fixed completion fixtures to compare includes with match location="exact" . One fixture was intentionally labelled as an incomplete-flag negative case. Our test expected the exact matcher to return incorrect, but Inspect returned C , causing this assertion failure: AssertionError: 'incomplete flag negative', 'exact', 'C', 'I' We do not treat that as proof that Inspect's scorer is defective. It shows that this fixture's expected result did not match the scorer's returned value. The assertion alone does not establish whether normalization, answer extraction, fixture construction, or another factor caused the mismatch. Because the assertion terminated the run, we did not obtain a complete five-fixture result set. That failure changed our implementation policy: every built-in scorer we adopt gets a table-driven contract suite containing punctuation, surrounding prose, denial text, case changes, whitespace, multiple answers, and truncated targets. A scorer name is not a specification. Partial credit introduces another interpretation problem. A score set of 0 , 0.5 , and 1 is easy to emit, but the business meaning of its aggregate is ours to define. A plain arithmetic mean treats two partial passes as equivalent to one full pass and one complete failure: mean 0.5, 0.5 == mean 1.0, 0.0 == 0.5 Those distributions carry different operational risk. We therefore preserve at least four outputs: full-pass rate, partial-pass rate, hard-failure rate, and mean score. We would not approve a model upgrade from the mean alone. We have not established byte-for-byte .eval reproducibility. Our offline scorer check did not run evaluations or compare log files; log reproducibility requires a separate experiment. Consequently, we have no defensible raw SHA-256 comparison to publish. Even after a successful run, we would separate two requirements: .eval file remains unchanged after creation. Raw byte equality is a stricter test. Run IDs, timestamps, ordering, provider metadata, and archive serialization can invalidate it even when the evaluation result is semantically identical. Our CI gate would hash the original artifact for custody, then generate a normalized comparison document for regression analysis. Model grading creates another source of nondeterminism. Exact match is brittle but inspectable. A model grader is flexible but adds a second model call, another prompt, another model version, and another failure mode. We did not complete a valid exact-match-versus-model-grader drift experiment, so we will not invent a disagreement rate. Our deployment rule is still clear: we pin the grader independently, store its explanation, and maintain a human-adjudicated calibration set. The sandbox boundary also needs precise language. Adding bash does not automatically isolate anything. In our task definition, sandbox="docker" is what assigns shell and Python execution to the container. We would not assume that custom Python tools are isolated; we would verify their execution context and route dangerous operations through a sandbox-aware interface. We therefore treat the evaluator host as sensitive infrastructure: Finally, we did not complete a provider rate-limit benchmark or a two-model token-cost comparison. The available run failed before remote inference. We cannot publish retry counts, wall time, throughput, or cost per sample from that attempt. The workaround is procedural, not rhetorical: record provider errors as their own failure category, pin concurrency, preserve token usage per sample, and rerun the exact same sample IDs against both model versions. We would reject any benchmark table that silently drops exhausted retries. Inspect's strongest comparison is not against an observability dashboard. It is against the bespoke evaluation runner that engineering teams eventually build around raw provider SDKs. | Option | Execution model | Item-level scoring | Agent trajectories | Isolation | Regression artifacts | Main cost | |---|---|---|---|---|---|---| | Inspect AI | Local Python runner with provider APIs | Strong | Strong | Explicit sandbox configuration | Structured .eval logs | Integration and evaluation design | | Bespoke Python harness | Whatever we implement | Variable | Usually custom work | Usually custom work | Usually JSONL or database rows | Engineering ownership | | Langfuse or Phoenix-style observability | Trace collection and analysis | Possible, but not the core runner | Strong tracing | Outside the primary scope | Trace-oriented | Backend operation and instrumentation | | Prompt-oriented CI harness | Configuration-driven test execution | Strong for prompt and security checks | Depends on integration | Depends on provider and tool setup | CI reports and artifacts | Configuration growth | | Commercial evaluation platform | Hosted execution and dashboards | Usually strong | Product-dependent | Product-dependent | Vendor-managed | Subscription, data governance, lock-in | We did not produce valid latency or throughput numbers, so we do not present fictional requests-per-second figures. For our deployment benchmark, we would measure provider latency, agent trajectory length, tool execution time, and local orchestration overhead separately before identifying the dominant bottleneck. Our cost gate uses artifact-derived quantities: sample cost = input tokens × input price per token + output tokens × output price per token + grader input tokens × grader input price per token + grader output tokens × grader output price per token + external tool cost We calculate the model-upgrade delta over identical sample IDs: full run delta = sum new model sample costs - sum old model sample costs We then combine quality and cost: cost per additional full pass = full run delta / new full pass count - old full pass count If the denominator is zero, this ratio is undefined; an upgrade can still improve cost efficiency by preserving quality at a lower full-run cost. Equal full-pass counts alone do not establish preserved quality, so we also review item-level regressions and failure categories. If the denominator is negative, we assess the quality loss and cost change separately rather than interpreting the ratio as a cost per additional full pass. If model grading is enabled, grader tokens stay separate from candidate-model tokens so that a grading prompt change cannot masquerade as candidate-model cost growth. The build-versus-buy calculation is similarly direct: annual internal harness cost = initial engineering hours × loaded hourly rate + annual maintenance hours × loaded hourly rate + incident and audit overhead annual inspect cost = integration hours × loaded hourly rate + task maintenance hours × loaded hourly rate + model and tool usage + sandbox infrastructure Inspect is open source, but it is not free to operate. The expensive parts are representative datasets, reliable scorers, model calls, sandbox execution, and review of ambiguous failures. Inspect removes a large amount of harness plumbing; it does not remove evaluation engineering. For broader implementation help, we keep our infrastructure patterns in the Effloow tools collection https://dev.to/tools . For teams turning an ad hoc benchmark into a release gate, our AI engineering services https://dev.to/services cover dataset versioning, grader calibration, and CI integration. We would deploy Inspect AI as the execution and artifact layer for a model-upgrade gate, with conditions. Deploy it if: Hold off or avoid it if: Our production verdict is adopt with explicit validation and operational controls . The component model is clean, the agent support is substantive, the task corpus is useful, and inspect view makes trajectory debugging materially easier. Inspect can replace weeks of basic harness construction. Our own validation was limited, however. Installation succeeded, but the offline scorer check terminated on an assertion mismatch before producing a complete five-fixture result set. Log reproducibility, provider rate-limit behavior, model-grading drift, and per-sample upgrade cost were outside that check's scope and remain untested. We would not promote that check into a production benchmark table. That restraint is part of the recommendation. Inspect provides the execution and logging components for an auditable release gate. We still have to prove the scorer, define the sandbox boundary, preserve the artifacts, and make the economics explicit. If those controls are already on your roadmap, Inspect is one of the first frameworks we would prototype. If you want help designing the gate around your own failure modes, contact Effloow https://dev.to/contact .