cd /news/ai-agents/inspect-ai-in-production-a-hard-nose… · home › topics › ai-agents › article
[ARTICLE · art-148960] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Inspect AI in Production: A Hard-Nosed Review of Agent Evals, Logs, and Model-Upgrade Gates

A developer evaluated Inspect AI as an execution and artifact layer for item-level agent-eval gates, concluding that it provides a stronger foundation than a home-grown request loop but leaves raw log reproducibility, scorer contracts, sandbox boundaries, and cost accounting as engineering responsibilities. The review, run with inspect-ai 0.3.100 on Python 3.12 in an isolated container, recommends pinning a small subset of the Inspect Evals corpus into an internal regression suite rather than importing a moving benchmark wholesale. The author notes that Inspect does not by itself make a benchmark deterministic, a model grader trustworthy, or a tool safe.

by read11 min views1 publishedOct 11, 2026

The model upgrade itself is rarely the difficult part. The difficult part is proving that the new model did not silently damage extraction accuracy, tool behavior, refusal handling, or operational cost.

Inspect AI Verdict

Inspect AI provides the execution and artifact layer for item-level agent-eval gates, but raw log reproducibility, scorer contracts, sandbox boundaries, and cost accounting remain engineering responsibilities.

We have seen the same internal harness emerge repeatedly: a JSONL dataset, an asynchronous request loop, a pile of retry logic, a scorer script, and a spreadsheet comparing the old model with the new one. It works until an incident forces us to answer questions the harness was never designed to answer:

We evaluated Inspect AI because it handles that entire execution layer. An Inspect evaluation combines a dataset, a solver, and one or more scorers. The solver may be a single generation call, a custom function decorated with @solver, or a multi-turn agent. The scorer may perform exact matching, inclusion matching, model grading, or arbitrary custom logic.

The framework also writes structured .eval logs and ships the inspect view browser interface. That separation matters. We do not want the dashboard to be the source of truth; we want the artifact to be the source of truth and the dashboard to be one way of reading it.

The agent surface was another reason for testing it. Inspect accepts agents in the solver position and includes a ReAct loop, tool interfaces, message and token limits, custom agents, agent handoffs, and adapters for external agents. Its tool catalog includes shell, Python, text editing, web, computer, and MCP-oriented interfaces. For dangerous execution, it supports sandbox backends including Docker and more infrastructure-heavy options.

We also inspected the Inspect Evals repository rather than evaluating the framework only with toy arithmetic. It provides a substantial task corpus and demonstrates how real evaluations organize datasets, task versions, dependencies, scoring, and sandbox assets. We still recommend pinning a small subset into an internal regression suite rather than importing a moving benchmark wholesale.

Our short conclusion before the details: Inspect gives us a much stronger foundation than another home-grown request loop. It does not, by itself, make a benchmark deterministic, a model grader trustworthy, or a tool safe.

We ran the offline scorer check with Python 3.12 in an isolated container. For the local walkthrough, we recommend a virtual environment. That was deliberate: Python 3.11 and 3.12 are the Inspect Evals project's preferred targets, while later Python versions carry compatibility caveats for parts of the broader task corpus.

For the local walkthrough, use this installation path:

python3.12 -m venv .venv
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install "inspect-ai==0.3.100" openai

export OPENAI_API_KEY="replace-me"
mkdir -p logs

The pinned inspect-ai wheel resolved successfully in our environment. The wheel we downloaded had this SHA-256 digest:

78e4c5b7e426fcc25563cd68f9f976a1c18c00adc9f4955bde9a6c97016363d7

For production, we would pin the complete lockfile rather than only the top-level package. Inspect pulls in a nontrivial dependency tree that includes asynchronous HTTP, schema, filesystem, terminal UI, and cloud-storage packages.

We then reduced the task shape to the smallest useful production example: registered task, registered solver, custom partial-credit scorer, stable sample IDs, and an explicit log directory.

import os
import re

from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.scorer import Score, Target, scorer
from inspect_ai.solver import Generate, TaskState, solver

@solver
def answer_once():
    async def solve(state: TaskState, generate: Generate) -> TaskState:
        return await generate(state)

    return solve

def normalize(text: str) -> str:
    return re.sub(r"\s+", " ", text.strip().lower())

@scorer
def partial_answer():
    async def score(state: TaskState, target: Target) -> Score:
        answer = normalize(state.output.completion)
        expected = normalize(target.text)

        if answer == expected:
            value = 1.0
            reason = "Exact normalized match"
        elif expected in answer:
            value = 0.5
            reason = "Target present with additional text"
        else:
            value = 0.0
            reason = "Target absent"

        return Score(
            value=value,
            answer=state.output.completion,
            explanation=reason,
        )

    return score

@task
def upgrade_gate():
    return Task(
        dataset=[
            Sample(
                id="refund-window",
                input="Reply with only the refund window: 30 days",
                target="30 days",
            ),
            Sample(
                id="support-tier",
                input="Reply with only the support tier: enterprise",
                target="enterprise",
            ),
        ],
        solver=answer_once(),
        scorer=partial_answer(),
    )

if __name__ == "__main__":
    eval(
        upgrade_gate(),
        model=os.environ.get("INSPECT_MODEL", "openai/gpt-5"),
        log_dir="logs",
    )

To execute this worked example through Python or the CLI and inspect its logs, use:

export INSPECT_MODEL="openai/gpt-5"
python regression_eval.py

inspect eval regression_eval.py@upgrade_gate \
  --model "$INSPECT_MODEL" \
  --log-dir logs

inspect view --log-dir logs

A successful run should produce output in this general shape. The following console block is illustrative, not a measurement from our failed qualification run:

$ inspect eval regression_eval.py@upgrade_gate \
    --model openai/gpt-5 \
    --log-dir logs

upgrade_gate (2 samples): complete
Model: openai/gpt-5
Log: logs/2026-10-11T101530_upgrade-gate.eval

Scores:
  partial_answer: 0.750

Samples:
  completed: 2
  errors:    0

For an agent task, the task definition changes more than the execution command. This is the minimal pattern we would use before adding a pinned challenge dataset:

from inspect_ai import Task, task
from inspect_ai.agent import react
from inspect_ai.scorer import includes
from inspect_ai.tool import bash, python, todo_write

@task
def sandbox_agent_eval():
    return Task(
        dataset=load_pinned_samples(),
        solver=react(
            prompt="Use the available tools and submit only the final answer.",
            tools=[bash(), python(), todo_write()],
            attempts=3,
        ),
        scorer=includes(),
        sandbox="docker",
        message_limit=30,
    )

The important production property is not the decorator syntax. It is that sample inputs, targets, messages, tool calls, outputs, scores, and run configuration can travel together in one reviewable artifact.

Our first narrow scorer contract test did not complete successfully.

We created five fixed completion fixtures to compare includes() with match(location="exact"). One fixture was intentionally labelled as an incomplete-flag negative case. Our test expected the exact matcher to return incorrect, but Inspect returned C, causing this assertion failure:

AssertionError:
('incomplete_flag_negative', 'exact', 'C', 'I')

We do not treat that as proof that Inspect's scorer is defective. It shows that this fixture's expected result did not match the scorer's returned value. The assertion alone does not establish whether normalization, answer extraction, fixture construction, or another factor caused the mismatch. Because the assertion terminated the run, we did not obtain a complete five-fixture result set.

That failure changed our implementation policy: every built-in scorer we adopt gets a table-driven contract suite containing punctuation, surrounding prose, denial text, case changes, whitespace, multiple answers, and truncated targets. A scorer name is not a specification.

Partial credit introduces another interpretation problem. A score set of 0, 0.5, and 1 is easy to emit, but the business meaning of its aggregate is ours to define. A plain arithmetic mean treats two partial passes as equivalent to one full pass and one complete failure:

mean([0.5, 0.5]) == mean([1.0, 0.0]) == 0.5

Those distributions carry different operational risk. We therefore preserve at least four outputs: full-pass rate, partial-pass rate, hard-failure rate, and mean score. We would not approve a model upgrade from the mean alone.

We have not established byte-for-byte .eval reproducibility. Our offline scorer check did not run evaluations or compare log files; log reproducibility requires a separate experiment. Consequently, we have no defensible raw SHA-256 comparison to publish.

Even after a successful run, we would separate two requirements:

.eval file remains unchanged after creation. Raw byte equality is a stricter test. Run IDs, timestamps, ordering, provider metadata, and archive serialization can invalidate it even when the evaluation result is semantically identical. Our CI gate would hash the original artifact for custody, then generate a normalized comparison document for regression analysis.

Model grading creates another source of nondeterminism. Exact match is brittle but inspectable. A model grader is flexible but adds a second model call, another prompt, another model version, and another failure mode. We did not complete a valid exact-match-versus-model-grader drift experiment, so we will not invent a disagreement rate. Our deployment rule is still clear: we pin the grader independently, store its explanation, and maintain a human-adjudicated calibration set.

The sandbox boundary also needs precise language. Adding bash() does not automatically isolate anything. In our task definition, sandbox="docker" is what assigns shell and Python execution to the container. We would not assume that custom Python tools are isolated; we would verify their execution context and route dangerous operations through a sandbox-aware interface.

We therefore treat the evaluator host as sensitive infrastructure:

Finally, we did not complete a provider rate-limit benchmark or a two-model token-cost comparison. The available run failed before remote inference. We cannot publish retry counts, wall time, throughput, or cost per sample from that attempt.

The workaround is procedural, not rhetorical: record provider errors as their own failure category, pin concurrency, preserve token usage per sample, and rerun the exact same sample IDs against both model versions. We would reject any benchmark table that silently drops exhausted retries.

Inspect's strongest comparison is not against an observability dashboard. It is against the bespoke evaluation runner that engineering teams eventually build around raw provider SDKs.

Option Execution model Item-level scoring Agent trajectories Isolation Regression artifacts Main cost
Inspect AI Local Python runner with provider APIs Strong Strong Explicit sandbox configuration Structured .eval logs Integration and evaluation design
Bespoke Python harness Whatever we implement Variable Usually custom work Usually custom work Usually JSONL or database rows Engineering ownership
Langfuse or Phoenix-style observability Trace collection and analysis Possible, but not the core runner Strong tracing Outside the primary scope Trace-oriented Backend operation and instrumentation
Prompt-oriented CI harness Configuration-driven test execution Strong for prompt and security checks Depends on integration Depends on provider and tool setup CI reports and artifacts Configuration growth
Commercial evaluation platform Hosted execution and dashboards Usually strong Product-dependent Product-dependent Vendor-managed Subscription, data governance, lock-in

We did not produce valid latency or throughput numbers, so we do not present fictional requests-per-second figures. For our deployment benchmark, we would measure provider latency, agent trajectory length, tool execution time, and local orchestration overhead separately before identifying the dominant bottleneck.

Our cost gate uses artifact-derived quantities:

sample_cost =
    input_tokens  × input_price_per_token
  + output_tokens × output_price_per_token
  + grader_input_tokens  × grader_input_price_per_token
  + grader_output_tokens × grader_output_price_per_token
  + external_tool_cost

We calculate the model-upgrade delta over identical sample IDs:

full_run_delta =
    sum(new_model_sample_costs)
  - sum(old_model_sample_costs)

We then combine quality and cost:

cost_per_additional_full_pass =
    full_run_delta
    / (new_full_pass_count - old_full_pass_count)

If the denominator is zero, this ratio is undefined; an upgrade can still improve cost efficiency by preserving quality at a lower full-run cost. Equal full-pass counts alone do not establish preserved quality, so we also review item-level regressions and failure categories. If the denominator is negative, we assess the quality loss and cost change separately rather than interpreting the ratio as a cost per additional full pass. If model grading is enabled, grader tokens stay separate from candidate-model tokens so that a grading prompt change cannot masquerade as candidate-model cost growth.

The build-versus-buy calculation is similarly direct:

annual_internal_harness_cost =
    initial_engineering_hours × loaded_hourly_rate
  + annual_maintenance_hours × loaded_hourly_rate
  + incident_and_audit_overhead

annual_inspect_cost =
    integration_hours × loaded_hourly_rate
  + task_maintenance_hours × loaded_hourly_rate
  + model_and_tool_usage
  + sandbox_infrastructure

Inspect is open source, but it is not free to operate. The expensive parts are representative datasets, reliable scorers, model calls, sandbox execution, and review of ambiguous failures. Inspect removes a large amount of harness plumbing; it does not remove evaluation engineering.

For broader implementation help, we keep our infrastructure patterns in the Effloow tools collection. For teams turning an ad hoc benchmark into a release gate, our AI engineering services cover dataset versioning, grader calibration, and CI integration.

We would deploy Inspect AI as the execution and artifact layer for a model-upgrade gate, with conditions.

Deploy it if:

Hold off or avoid it if:

Our production verdict is adopt with explicit validation and operational controls.

The component model is clean, the agent support is substantive, the task corpus is useful, and inspect view makes trajectory debugging materially easier. Inspect can replace weeks of basic harness construction.

Our own validation was limited, however. Installation succeeded, but the offline scorer check terminated on an assertion mismatch before producing a complete five-fixture result set. Log reproducibility, provider rate-limit behavior, model-grading drift, and per-sample upgrade cost were outside that check's scope and remain untested. We would not promote that check into a production benchmark table.

That restraint is part of the recommendation. Inspect provides the execution and logging components for an auditable release gate. We still have to prove the scorer, define the sandbox boundary, preserve the artifacts, and make the economics explicit.

If those controls are already on your roadmap, Inspect is one of the first frameworks we would prototype. If you want help designing the gate around your own failure modes, contact Effloow.

── more in #ai-agents 4 stories · sorted by recency
── more on @inspect ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/inspect-ai-in-produc…] indexed:0 read:11min 2026-10-11 · —