# Inspect AI in Production: A Hard-Nosed Review of Agent Evals, Logs, and Model-Upgrade Gates

> Source: <https://dev.to/jangwook_kim_e31e7291ad98/inspect-ai-in-production-a-hard-nosed-review-of-agent-evals-logs-and-model-upgrade-gates-1mj5>
> Published: 2026-10-11 00:39:48+00:00

The model upgrade itself is rarely the difficult part. The difficult part is proving that the new model did not silently damage extraction accuracy, tool behavior, refusal handling, or operational cost.

Inspect AI Verdict

Inspect AI provides the execution and artifact layer for item-level agent-eval gates, but raw log reproducibility, scorer contracts, sandbox boundaries, and cost accounting remain engineering responsibilities.

We have seen the same internal harness emerge repeatedly: a JSONL dataset, an asynchronous request loop, a pile of retry logic, a scorer script, and a spreadsheet comparing the old model with the new one. It works until an incident forces us to answer questions the harness was never designed to answer:

We evaluated [Inspect AI](https://inspect.aisi.org.uk/) because it handles that entire execution layer. An Inspect evaluation combines a dataset, a solver, and one or more scorers. The solver may be a single generation call, a custom function decorated with `@solver`, or a multi-turn agent. The scorer may perform exact matching, inclusion matching, model grading, or arbitrary custom logic.

The framework also writes structured `.eval` logs and ships the `inspect view` browser interface. That separation matters. We do not want the dashboard to be the source of truth; we want the artifact to be the source of truth and the dashboard to be one way of reading it.

The agent surface was another reason for testing it. Inspect accepts agents in the solver position and includes a ReAct loop, tool interfaces, message and token limits, custom agents, agent handoffs, and adapters for external agents. Its tool catalog includes shell, Python, text editing, web, computer, and MCP-oriented interfaces. For dangerous execution, it supports sandbox backends including Docker and more infrastructure-heavy options.

We also inspected the [Inspect Evals repository](https://github.com/UKGovernmentBEIS/inspect_evals) rather than evaluating the framework only with toy arithmetic. It provides a substantial task corpus and demonstrates how real evaluations organize datasets, task versions, dependencies, scoring, and sandbox assets. We still recommend pinning a small subset into an internal regression suite rather than importing a moving benchmark wholesale.

Our short conclusion before the details: Inspect gives us a much stronger foundation than another home-grown request loop. It does not, by itself, make a benchmark deterministic, a model grader trustworthy, or a tool safe.

We ran the offline scorer check with Python 3.12 in an isolated container. For the local walkthrough, we recommend a virtual environment. That was deliberate: Python 3.11 and 3.12 are the Inspect Evals project's preferred targets, while later Python versions carry compatibility caveats for parts of the broader task corpus.

For the local walkthrough, use this installation path:

```
python3.12 -m venv .venv
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install "inspect-ai==0.3.100" openai

export OPENAI_API_KEY="replace-me"
mkdir -p logs
```

The pinned `inspect-ai` wheel resolved successfully in our environment. The wheel we downloaded had this SHA-256 digest:

```
78e4c5b7e426fcc25563cd68f9f976a1c18c00adc9f4955bde9a6c97016363d7
```

For production, we would pin the complete lockfile rather than only the top-level package. Inspect pulls in a nontrivial dependency tree that includes asynchronous HTTP, schema, filesystem, terminal UI, and cloud-storage packages.

We then reduced the task shape to the smallest useful production example: registered task, registered solver, custom partial-credit scorer, stable sample IDs, and an explicit log directory.

``` python
# regression_eval.py
import os
import re

from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.scorer import Score, Target, scorer
from inspect_ai.solver import Generate, TaskState, solver

@solver
def answer_once():
    async def solve(state: TaskState, generate: Generate) -> TaskState:
        return await generate(state)

    return solve

def normalize(text: str) -> str:
    return re.sub(r"\s+", " ", text.strip().lower())

@scorer
def partial_answer():
    async def score(state: TaskState, target: Target) -> Score:
        answer = normalize(state.output.completion)
        expected = normalize(target.text)

        if answer == expected:
            value = 1.0
            reason = "Exact normalized match"
        elif expected in answer:
            value = 0.5
            reason = "Target present with additional text"
        else:
            value = 0.0
            reason = "Target absent"

        return Score(
            value=value,
            answer=state.output.completion,
            explanation=reason,
        )

    return score

@task
def upgrade_gate():
    return Task(
        dataset=[
            Sample(
                id="refund-window",
                input="Reply with only the refund window: 30 days",
                target="30 days",
            ),
            Sample(
                id="support-tier",
                input="Reply with only the support tier: enterprise",
                target="enterprise",
            ),
        ],
        solver=answer_once(),
        scorer=partial_answer(),
    )

if __name__ == "__main__":
    eval(
        upgrade_gate(),
        model=os.environ.get("INSPECT_MODEL", "openai/gpt-5"),
        log_dir="logs",
    )
```

To execute this worked example through Python or the CLI and inspect its logs, use:

```
export INSPECT_MODEL="openai/gpt-5"
python regression_eval.py

# Equivalent task discovery through the CLI
inspect eval regression_eval.py@upgrade_gate \
  --model "$INSPECT_MODEL" \
  --log-dir logs

inspect view --log-dir logs
```

A successful run should produce output in this general shape. The following console block is illustrative, not a measurement from our failed qualification run:

``` bash
$ inspect eval regression_eval.py@upgrade_gate \
    --model openai/gpt-5 \
    --log-dir logs

upgrade_gate (2 samples): complete
Model: openai/gpt-5
Log: logs/2026-10-11T101530_upgrade-gate.eval

Scores:
  partial_answer: 0.750

Samples:
  completed: 2
  errors:    0
```

For an agent task, the task definition changes more than the execution command. This is the minimal pattern we would use before adding a pinned challenge dataset:

``` python
from inspect_ai import Task, task
from inspect_ai.agent import react
from inspect_ai.scorer import includes
from inspect_ai.tool import bash, python, todo_write

@task
def sandbox_agent_eval():
    return Task(
        dataset=load_pinned_samples(),
        solver=react(
            prompt="Use the available tools and submit only the final answer.",
            tools=[bash(), python(), todo_write()],
            attempts=3,
        ),
        scorer=includes(),
        sandbox="docker",
        message_limit=30,
    )
```

The important production property is not the decorator syntax. It is that sample inputs, targets, messages, tool calls, outputs, scores, and run configuration can travel together in one reviewable artifact.

Our first narrow scorer contract test did not complete successfully.

We created five fixed completion fixtures to compare `includes()` with `match(location="exact")`. One fixture was intentionally labelled as an incomplete-flag negative case. Our test expected the exact matcher to return incorrect, but Inspect returned `C`, causing this assertion failure:

```
AssertionError:
('incomplete_flag_negative', 'exact', 'C', 'I')
```

We do not treat that as proof that Inspect's scorer is defective. It shows that this fixture's expected result did not match the scorer's returned value. The assertion alone does not establish whether normalization, answer extraction, fixture construction, or another factor caused the mismatch. Because the assertion terminated the run, we did not obtain a complete five-fixture result set.

That failure changed our implementation policy: every built-in scorer we adopt gets a table-driven contract suite containing punctuation, surrounding prose, denial text, case changes, whitespace, multiple answers, and truncated targets. A scorer name is not a specification.

Partial credit introduces another interpretation problem. A score set of `0`, `0.5`, and `1` is easy to emit, but the business meaning of its aggregate is ours to define. A plain arithmetic mean treats two partial passes as equivalent to one full pass and one complete failure:

```
mean([0.5, 0.5]) == mean([1.0, 0.0]) == 0.5
```

Those distributions carry different operational risk. We therefore preserve at least four outputs: full-pass rate, partial-pass rate, hard-failure rate, and mean score. We would not approve a model upgrade from the mean alone.

We have not established byte-for-byte `.eval` reproducibility. Our offline scorer check did not run evaluations or compare log files; log reproducibility requires a separate experiment. Consequently, we have no defensible raw SHA-256 comparison to publish.

Even after a successful run, we would separate two requirements:

`.eval` file remains unchanged after creation.
Raw byte equality is a stricter test. Run IDs, timestamps, ordering, provider metadata, and archive serialization can invalidate it even when the evaluation result is semantically identical. Our CI gate would hash the original artifact for custody, then generate a normalized comparison document for regression analysis.

Model grading creates another source of nondeterminism. Exact match is brittle but inspectable. A model grader is flexible but adds a second model call, another prompt, another model version, and another failure mode. We did not complete a valid exact-match-versus-model-grader drift experiment, so we will not invent a disagreement rate. Our deployment rule is still clear: we pin the grader independently, store its explanation, and maintain a human-adjudicated calibration set.

The sandbox boundary also needs precise language. Adding `bash()` does not automatically isolate anything. In our task definition, `sandbox="docker"` is what assigns shell and Python execution to the container. We would not assume that custom Python tools are isolated; we would verify their execution context and route dangerous operations through a sandbox-aware interface.

We therefore treat the evaluator host as sensitive infrastructure:

Finally, we did not complete a provider rate-limit benchmark or a two-model token-cost comparison. The available run failed before remote inference. We cannot publish retry counts, wall time, throughput, or cost per sample from that attempt.

The workaround is procedural, not rhetorical: record provider errors as their own failure category, pin concurrency, preserve token usage per sample, and rerun the exact same sample IDs against both model versions. We would reject any benchmark table that silently drops exhausted retries.

Inspect's strongest comparison is not against an observability dashboard. It is against the bespoke evaluation runner that engineering teams eventually build around raw provider SDKs.

| Option | Execution model | Item-level scoring | Agent trajectories | Isolation | Regression artifacts | Main cost | 
|---|---|---|---|---|---|---|
| Inspect AI | Local Python runner with provider APIs | Strong | Strong | Explicit sandbox configuration | Structured `.eval` logs | Integration and evaluation design | 
| Bespoke Python harness | Whatever we implement | Variable | Usually custom work | Usually custom work | Usually JSONL or database rows | Engineering ownership | 
| Langfuse or Phoenix-style observability | Trace collection and analysis | Possible, but not the core runner | Strong tracing | Outside the primary scope | Trace-oriented | Backend operation and instrumentation | 
| Prompt-oriented CI harness | Configuration-driven test execution | Strong for prompt and security checks | Depends on integration | Depends on provider and tool setup | CI reports and artifacts | Configuration growth | 
| Commercial evaluation platform | Hosted execution and dashboards | Usually strong | Product-dependent | Product-dependent | Vendor-managed | Subscription, data governance, lock-in | 

We did not produce valid latency or throughput numbers, so we do not present fictional requests-per-second figures. For our deployment benchmark, we would measure provider latency, agent trajectory length, tool execution time, and local orchestration overhead separately before identifying the dominant bottleneck.

Our cost gate uses artifact-derived quantities:

```
sample_cost =
    input_tokens  × input_price_per_token
  + output_tokens × output_price_per_token
  + grader_input_tokens  × grader_input_price_per_token
  + grader_output_tokens × grader_output_price_per_token
  + external_tool_cost
```

We calculate the model-upgrade delta over identical sample IDs:

```
full_run_delta =
    sum(new_model_sample_costs)
  - sum(old_model_sample_costs)
```

We then combine quality and cost:

```
cost_per_additional_full_pass =
    full_run_delta
    / (new_full_pass_count - old_full_pass_count)
```

If the denominator is zero, this ratio is undefined; an upgrade can still improve cost efficiency by preserving quality at a lower full-run cost. Equal full-pass counts alone do not establish preserved quality, so we also review item-level regressions and failure categories. If the denominator is negative, we assess the quality loss and cost change separately rather than interpreting the ratio as a cost per additional full pass. If model grading is enabled, grader tokens stay separate from candidate-model tokens so that a grading prompt change cannot masquerade as candidate-model cost growth.

The build-versus-buy calculation is similarly direct:

```
annual_internal_harness_cost =
    initial_engineering_hours × loaded_hourly_rate
  + annual_maintenance_hours × loaded_hourly_rate
  + incident_and_audit_overhead

annual_inspect_cost =
    integration_hours × loaded_hourly_rate
  + task_maintenance_hours × loaded_hourly_rate
  + model_and_tool_usage
  + sandbox_infrastructure
```

Inspect is open source, but it is not free to operate. The expensive parts are representative datasets, reliable scorers, model calls, sandbox execution, and review of ambiguous failures. Inspect removes a large amount of harness plumbing; it does not remove evaluation engineering.

For broader implementation help, we keep our infrastructure patterns in the [Effloow tools collection](https://dev.to/tools). For teams turning an ad hoc benchmark into a release gate, our [AI engineering services](https://dev.to/services) cover dataset versioning, grader calibration, and CI integration.

We would deploy Inspect AI as the execution and artifact layer for a model-upgrade gate, with conditions.

**Deploy it if:**

**Hold off or avoid it if:**

Our production verdict is **adopt with explicit validation and operational controls**.

The component model is clean, the agent support is substantive, the task corpus is useful, and `inspect view` makes trajectory debugging materially easier. Inspect can replace weeks of basic harness construction.

Our own validation was limited, however. Installation succeeded, but the offline scorer check terminated on an assertion mismatch before producing a complete five-fixture result set. Log reproducibility, provider rate-limit behavior, model-grading drift, and per-sample upgrade cost were outside that check's scope and remain untested. We would not promote that check into a production benchmark table.

That restraint is part of the recommendation. Inspect provides the execution and logging components for an auditable release gate. We still have to prove the scorer, define the sandbox boundary, preserve the artifacts, and make the economics explicit.

If those controls are already on your roadmap, Inspect is one of the first frameworks we would prototype. If you want help designing the gate around your own failure modes, [contact Effloow](https://dev.to/contact).
