{"slug": "inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-gates", "title": "Inspect AI in Production: A Hard-Nosed Review of Agent Evals, Logs, and Model-Upgrade Gates", "summary": "A developer evaluated Inspect AI as an execution and artifact layer for item-level agent-eval gates, concluding that it provides a stronger foundation than a home-grown request loop but leaves raw log reproducibility, scorer contracts, sandbox boundaries, and cost accounting as engineering responsibilities. The review, run with inspect-ai 0.3.100 on Python 3.12 in an isolated container, recommends pinning a small subset of the Inspect Evals corpus into an internal regression suite rather than importing a moving benchmark wholesale. The author notes that Inspect does not by itself make a benchmark deterministic, a model grader trustworthy, or a tool safe.", "body_md": "The model upgrade itself is rarely the difficult part. The difficult part is proving that the new model did not silently damage extraction accuracy, tool behavior, refusal handling, or operational cost.\n\nInspect AI Verdict\n\nInspect AI provides the execution and artifact layer for item-level agent-eval gates, but raw log reproducibility, scorer contracts, sandbox boundaries, and cost accounting remain engineering responsibilities.\n\nWe have seen the same internal harness emerge repeatedly: a JSONL dataset, an asynchronous request loop, a pile of retry logic, a scorer script, and a spreadsheet comparing the old model with the new one. It works until an incident forces us to answer questions the harness was never designed to answer:\n\nWe evaluated [Inspect AI](https://inspect.aisi.org.uk/) because it handles that entire execution layer. An Inspect evaluation combines a dataset, a solver, and one or more scorers. The solver may be a single generation call, a custom function decorated with `@solver`, or a multi-turn agent. The scorer may perform exact matching, inclusion matching, model grading, or arbitrary custom logic.\n\nThe framework also writes structured `.eval` logs and ships the `inspect view` browser interface. That separation matters. We do not want the dashboard to be the source of truth; we want the artifact to be the source of truth and the dashboard to be one way of reading it.\n\nThe agent surface was another reason for testing it. Inspect accepts agents in the solver position and includes a ReAct loop, tool interfaces, message and token limits, custom agents, agent handoffs, and adapters for external agents. Its tool catalog includes shell, Python, text editing, web, computer, and MCP-oriented interfaces. For dangerous execution, it supports sandbox backends including Docker and more infrastructure-heavy options.\n\nWe also inspected the [Inspect Evals repository](https://github.com/UKGovernmentBEIS/inspect_evals) rather than evaluating the framework only with toy arithmetic. It provides a substantial task corpus and demonstrates how real evaluations organize datasets, task versions, dependencies, scoring, and sandbox assets. We still recommend pinning a small subset into an internal regression suite rather than importing a moving benchmark wholesale.\n\nOur short conclusion before the details: Inspect gives us a much stronger foundation than another home-grown request loop. It does not, by itself, make a benchmark deterministic, a model grader trustworthy, or a tool safe.\n\nWe ran the offline scorer check with Python 3.12 in an isolated container. For the local walkthrough, we recommend a virtual environment. That was deliberate: Python 3.11 and 3.12 are the Inspect Evals project's preferred targets, while later Python versions carry compatibility caveats for parts of the broader task corpus.\n\nFor the local walkthrough, use this installation path:\n\n```\npython3.12 -m venv .venv\nsource .venv/bin/activate\n\npython -m pip install --upgrade pip\npython -m pip install \"inspect-ai==0.3.100\" openai\n\nexport OPENAI_API_KEY=\"replace-me\"\nmkdir -p logs\n```\n\nThe pinned `inspect-ai` wheel resolved successfully in our environment. The wheel we downloaded had this SHA-256 digest:\n\n```\n78e4c5b7e426fcc25563cd68f9f976a1c18c00adc9f4955bde9a6c97016363d7\n```\n\nFor production, we would pin the complete lockfile rather than only the top-level package. Inspect pulls in a nontrivial dependency tree that includes asynchronous HTTP, schema, filesystem, terminal UI, and cloud-storage packages.\n\nWe then reduced the task shape to the smallest useful production example: registered task, registered solver, custom partial-credit scorer, stable sample IDs, and an explicit log directory.\n\n``` python\n# regression_eval.py\nimport os\nimport re\n\nfrom inspect_ai import Task, eval, task\nfrom inspect_ai.dataset import Sample\nfrom inspect_ai.scorer import Score, Target, scorer\nfrom inspect_ai.solver import Generate, TaskState, solver\n\n@solver\ndef answer_once():\n    async def solve(state: TaskState, generate: Generate) -> TaskState:\n        return await generate(state)\n\n    return solve\n\ndef normalize(text: str) -> str:\n    return re.sub(r\"\\s+\", \" \", text.strip().lower())\n\n@scorer\ndef partial_answer():\n    async def score(state: TaskState, target: Target) -> Score:\n        answer = normalize(state.output.completion)\n        expected = normalize(target.text)\n\n        if answer == expected:\n            value = 1.0\n            reason = \"Exact normalized match\"\n        elif expected in answer:\n            value = 0.5\n            reason = \"Target present with additional text\"\n        else:\n            value = 0.0\n            reason = \"Target absent\"\n\n        return Score(\n            value=value,\n            answer=state.output.completion,\n            explanation=reason,\n        )\n\n    return score\n\n@task\ndef upgrade_gate():\n    return Task(\n        dataset=[\n            Sample(\n                id=\"refund-window\",\n                input=\"Reply with only the refund window: 30 days\",\n                target=\"30 days\",\n            ),\n            Sample(\n                id=\"support-tier\",\n                input=\"Reply with only the support tier: enterprise\",\n                target=\"enterprise\",\n            ),\n        ],\n        solver=answer_once(),\n        scorer=partial_answer(),\n    )\n\nif __name__ == \"__main__\":\n    eval(\n        upgrade_gate(),\n        model=os.environ.get(\"INSPECT_MODEL\", \"openai/gpt-5\"),\n        log_dir=\"logs\",\n    )\n```\n\nTo execute this worked example through Python or the CLI and inspect its logs, use:\n\n```\nexport INSPECT_MODEL=\"openai/gpt-5\"\npython regression_eval.py\n\n# Equivalent task discovery through the CLI\ninspect eval regression_eval.py@upgrade_gate \\\n  --model \"$INSPECT_MODEL\" \\\n  --log-dir logs\n\ninspect view --log-dir logs\n```\n\nA successful run should produce output in this general shape. The following console block is illustrative, not a measurement from our failed qualification run:\n\n``` bash\n$ inspect eval regression_eval.py@upgrade_gate \\\n    --model openai/gpt-5 \\\n    --log-dir logs\n\nupgrade_gate (2 samples): complete\nModel: openai/gpt-5\nLog: logs/2026-10-11T101530_upgrade-gate.eval\n\nScores:\n  partial_answer: 0.750\n\nSamples:\n  completed: 2\n  errors:    0\n```\n\nFor an agent task, the task definition changes more than the execution command. This is the minimal pattern we would use before adding a pinned challenge dataset:\n\n``` python\nfrom inspect_ai import Task, task\nfrom inspect_ai.agent import react\nfrom inspect_ai.scorer import includes\nfrom inspect_ai.tool import bash, python, todo_write\n\n@task\ndef sandbox_agent_eval():\n    return Task(\n        dataset=load_pinned_samples(),\n        solver=react(\n            prompt=\"Use the available tools and submit only the final answer.\",\n            tools=[bash(), python(), todo_write()],\n            attempts=3,\n        ),\n        scorer=includes(),\n        sandbox=\"docker\",\n        message_limit=30,\n    )\n```\n\nThe important production property is not the decorator syntax. It is that sample inputs, targets, messages, tool calls, outputs, scores, and run configuration can travel together in one reviewable artifact.\n\nOur first narrow scorer contract test did not complete successfully.\n\nWe created five fixed completion fixtures to compare `includes()` with `match(location=\"exact\")`. One fixture was intentionally labelled as an incomplete-flag negative case. Our test expected the exact matcher to return incorrect, but Inspect returned `C`, causing this assertion failure:\n\n```\nAssertionError:\n('incomplete_flag_negative', 'exact', 'C', 'I')\n```\n\nWe do not treat that as proof that Inspect's scorer is defective. It shows that this fixture's expected result did not match the scorer's returned value. The assertion alone does not establish whether normalization, answer extraction, fixture construction, or another factor caused the mismatch. Because the assertion terminated the run, we did not obtain a complete five-fixture result set.\n\nThat failure changed our implementation policy: every built-in scorer we adopt gets a table-driven contract suite containing punctuation, surrounding prose, denial text, case changes, whitespace, multiple answers, and truncated targets. A scorer name is not a specification.\n\nPartial credit introduces another interpretation problem. A score set of `0`, `0.5`, and `1` is easy to emit, but the business meaning of its aggregate is ours to define. A plain arithmetic mean treats two partial passes as equivalent to one full pass and one complete failure:\n\n```\nmean([0.5, 0.5]) == mean([1.0, 0.0]) == 0.5\n```\n\nThose distributions carry different operational risk. We therefore preserve at least four outputs: full-pass rate, partial-pass rate, hard-failure rate, and mean score. We would not approve a model upgrade from the mean alone.\n\nWe have not established byte-for-byte `.eval` reproducibility. Our offline scorer check did not run evaluations or compare log files; log reproducibility requires a separate experiment. Consequently, we have no defensible raw SHA-256 comparison to publish.\n\nEven after a successful run, we would separate two requirements:\n\n`.eval` file remains unchanged after creation.\nRaw byte equality is a stricter test. Run IDs, timestamps, ordering, provider metadata, and archive serialization can invalidate it even when the evaluation result is semantically identical. Our CI gate would hash the original artifact for custody, then generate a normalized comparison document for regression analysis.\n\nModel grading creates another source of nondeterminism. Exact match is brittle but inspectable. A model grader is flexible but adds a second model call, another prompt, another model version, and another failure mode. We did not complete a valid exact-match-versus-model-grader drift experiment, so we will not invent a disagreement rate. Our deployment rule is still clear: we pin the grader independently, store its explanation, and maintain a human-adjudicated calibration set.\n\nThe sandbox boundary also needs precise language. Adding `bash()` does not automatically isolate anything. In our task definition, `sandbox=\"docker\"` is what assigns shell and Python execution to the container. We would not assume that custom Python tools are isolated; we would verify their execution context and route dangerous operations through a sandbox-aware interface.\n\nWe therefore treat the evaluator host as sensitive infrastructure:\n\nFinally, we did not complete a provider rate-limit benchmark or a two-model token-cost comparison. The available run failed before remote inference. We cannot publish retry counts, wall time, throughput, or cost per sample from that attempt.\n\nThe workaround is procedural, not rhetorical: record provider errors as their own failure category, pin concurrency, preserve token usage per sample, and rerun the exact same sample IDs against both model versions. We would reject any benchmark table that silently drops exhausted retries.\n\nInspect's strongest comparison is not against an observability dashboard. It is against the bespoke evaluation runner that engineering teams eventually build around raw provider SDKs.\n\n| Option | Execution model | Item-level scoring | Agent trajectories | Isolation | Regression artifacts | Main cost | \n|---|---|---|---|---|---|---|\n| Inspect AI | Local Python runner with provider APIs | Strong | Strong | Explicit sandbox configuration | Structured `.eval` logs | Integration and evaluation design | \n| Bespoke Python harness | Whatever we implement | Variable | Usually custom work | Usually custom work | Usually JSONL or database rows | Engineering ownership | \n| Langfuse or Phoenix-style observability | Trace collection and analysis | Possible, but not the core runner | Strong tracing | Outside the primary scope | Trace-oriented | Backend operation and instrumentation | \n| Prompt-oriented CI harness | Configuration-driven test execution | Strong for prompt and security checks | Depends on integration | Depends on provider and tool setup | CI reports and artifacts | Configuration growth | \n| Commercial evaluation platform | Hosted execution and dashboards | Usually strong | Product-dependent | Product-dependent | Vendor-managed | Subscription, data governance, lock-in | \n\nWe did not produce valid latency or throughput numbers, so we do not present fictional requests-per-second figures. For our deployment benchmark, we would measure provider latency, agent trajectory length, tool execution time, and local orchestration overhead separately before identifying the dominant bottleneck.\n\nOur cost gate uses artifact-derived quantities:\n\n```\nsample_cost =\n    input_tokens  × input_price_per_token\n  + output_tokens × output_price_per_token\n  + grader_input_tokens  × grader_input_price_per_token\n  + grader_output_tokens × grader_output_price_per_token\n  + external_tool_cost\n```\n\nWe calculate the model-upgrade delta over identical sample IDs:\n\n```\nfull_run_delta =\n    sum(new_model_sample_costs)\n  - sum(old_model_sample_costs)\n```\n\nWe then combine quality and cost:\n\n```\ncost_per_additional_full_pass =\n    full_run_delta\n    / (new_full_pass_count - old_full_pass_count)\n```\n\nIf the denominator is zero, this ratio is undefined; an upgrade can still improve cost efficiency by preserving quality at a lower full-run cost. Equal full-pass counts alone do not establish preserved quality, so we also review item-level regressions and failure categories. If the denominator is negative, we assess the quality loss and cost change separately rather than interpreting the ratio as a cost per additional full pass. If model grading is enabled, grader tokens stay separate from candidate-model tokens so that a grading prompt change cannot masquerade as candidate-model cost growth.\n\nThe build-versus-buy calculation is similarly direct:\n\n```\nannual_internal_harness_cost =\n    initial_engineering_hours × loaded_hourly_rate\n  + annual_maintenance_hours × loaded_hourly_rate\n  + incident_and_audit_overhead\n\nannual_inspect_cost =\n    integration_hours × loaded_hourly_rate\n  + task_maintenance_hours × loaded_hourly_rate\n  + model_and_tool_usage\n  + sandbox_infrastructure\n```\n\nInspect is open source, but it is not free to operate. The expensive parts are representative datasets, reliable scorers, model calls, sandbox execution, and review of ambiguous failures. Inspect removes a large amount of harness plumbing; it does not remove evaluation engineering.\n\nFor broader implementation help, we keep our infrastructure patterns in the [Effloow tools collection](https://dev.to/tools). For teams turning an ad hoc benchmark into a release gate, our [AI engineering services](https://dev.to/services) cover dataset versioning, grader calibration, and CI integration.\n\nWe would deploy Inspect AI as the execution and artifact layer for a model-upgrade gate, with conditions.\n\n**Deploy it if:**\n\n**Hold off or avoid it if:**\n\nOur production verdict is **adopt with explicit validation and operational controls**.\n\nThe component model is clean, the agent support is substantive, the task corpus is useful, and `inspect view` makes trajectory debugging materially easier. Inspect can replace weeks of basic harness construction.\n\nOur own validation was limited, however. Installation succeeded, but the offline scorer check terminated on an assertion mismatch before producing a complete five-fixture result set. Log reproducibility, provider rate-limit behavior, model-grading drift, and per-sample upgrade cost were outside that check's scope and remain untested. We would not promote that check into a production benchmark table.\n\nThat restraint is part of the recommendation. Inspect provides the execution and logging components for an auditable release gate. We still have to prove the scorer, define the sandbox boundary, preserve the artifacts, and make the economics explicit.\n\nIf those controls are already on your roadmap, Inspect is one of the first frameworks we would prototype. If you want help designing the gate around your own failure modes, [contact Effloow](https://dev.to/contact).", "url": "https://wpnews.pro/news/inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-gates", "canonical_source": "https://dev.to/jangwook_kim_e31e7291ad98/inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-upgrade-gates-1mj5", "published_at": "2026-10-11 00:39:48+00:00", "updated_at": "2026-10-11 00:49:50.808257+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "mlops", "developer-tools", "ai-safety"], "entities": ["Inspect AI", "Inspect Evals", "UKGovernmentBEIS", "OpenAI", "Docker", "Python", "MCP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-gates", "markdown": "https://wpnews.pro/news/inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-gates.md", "text": "https://wpnews.pro/news/inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-gates.txt", "jsonld": "https://wpnews.pro/news/inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-gates.jsonld"}}