cd /news/ai-agents/openai-and-ironclad-turning-contract… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-146443] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=↑ positive

OpenAI and Ironclad: Turning Contract Workflows Into Agent Evals

OpenAI and Ironclad published a case study describing how Ironclad's production contract lifecycle workflows are being converted into reproducible evaluation benchmarks for computer-use agents. The approach sanitizes real multi-step approval, redlining and negotiation flows into synthetic contracts, versions eval snapshots against product releases, and uses human-labeled failure taxonomies to retrain the agents.

by read5 min views1 publishedOct 7, 2026

OpenAI and Ironclad just published a case study on using production contract workflows as both training data and evaluation benchmarks for computer-use agents. This is not a demo. It is a partnership where a SaaS company opens its workflow engine to become an agent training ground.

The plumbing question is simple: how do you turn a multi-step contract approval flow into a reproducible eval without leaking customer data, and how do you keep that eval valid when the underlying product changes?

Most computer-use benchmarks are synthetic. They simulate browser tasks or API calls in controlled environments. Ironclad's contract lifecycle management platform offers something different: real multi-step workflows with approval chains, redlining, negotiation loops, and conditional branching.

These workflows are stateful. A contract might move from draft to legal review, back to sales for edits, then to finance for approval. Each step has different permissions, different UI surfaces, and different failure modes. If an agent can navigate this, it can navigate most SaaS tools.

The partnership gives OpenAI access to:

Turning a production workflow into an agent eval requires solving three problems:

1. Data sanitization

You cannot train on customer contracts. Ironclad must generate synthetic contracts that preserve workflow complexity without exposing real terms, parties, or negotiation history. This means:

2. Reproducibility

A contract workflow is not deterministic. Different users take different paths. To make this an eval, you need:

3. Version drift

Ironclad ships product updates. When the UI changes, the eval breaks. The infrastructure must:

Here is the likely flow:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Ironclad Prod   β”‚
β”‚ Workflow Engine β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Workflow Logger β”‚  ← Captures state transitions, UI events, API calls
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Sanitization    β”‚  ← Strips PII, generates synthetic contracts
β”‚ Pipeline        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Eval Snapshot   β”‚  ← Versioned environment + workflow definition
β”‚ Generator       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Agent Executor  β”‚  ← OpenAI agent runs against snapshot
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Failure Labeler β”‚  ← Human annotators classify agent errors
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The workflow logger is the critical piece. It must capture:

When an agent fails on a contract task, someone has to classify why. This is not automated. The labeling taxonomy likely includes:

Failure Type Example Fix Strategy
Navigation error Clicked wrong button Improve UI element detection
State misread Thought contract was approved when pending Better state extraction from DOM
Logic error Skipped required approval step Refine workflow graph understanding
Timeout Took too long on redline comparison Optimize document parsing
Permission boundary Tried to approve without authority Teach role-based access model

Ironclad employees likely do the initial labeling. OpenAI uses those labels to retrain. The loop tightens over time.

A reproducible eval needs a snapshot format. Here is a plausible schema:

from dataclasses import dataclass
from typing import List, Dict, Any

@dataclass
class WorkflowSnapshot:
    version: str  # Ironclad product version
    workflow_id: str
    initial_state: Dict[str, Any]  # Contract metadata, approvers, etc.
    steps: List[Dict[str, Any]]  # Ordered list of expected actions
    success_criteria: Dict[str, Any]  # What "done" looks like
    ui_snapshots: List[str]  # DOM or screenshot hashes per step

    def validate_agent_run(self, agent_trace: List[Dict]) -> bool:
        """
        Compare agent actions against expected workflow steps.
        Returns True if agent reached success criteria.
        """
        for i, expected_step in enumerate(self.steps):
            if i >= len(agent_trace):
                return False  # Agent stopped early

            actual = agent_trace[i]
            if not self._step_matches(expected_step, actual):
                return False

        return self._check_success(agent_trace[-1])

    def _step_matches(self, expected: Dict, actual: Dict) -> bool:
        return (
            expected["action_type"] == actual["action_type"] and
            expected["target_element"] in actual["dom_path"]
        )

    def _check_success(self, final_state: Dict) -> bool:
        return final_state.get("contract_status") == self.success_criteria.get("status")

This snapshot is versioned. When Ironclad ships a UI update, old snapshots remain valid for regression testing. New snapshots get generated for the updated product.

To debug agent failures, you need full trace data:

The observability stack must correlate these streams. If an agent clicks the wrong button, you need to see:

Without this correlation, failure labeling becomes guesswork.

When Ironclad changes a workflow (new approval step, different UI layout), the eval suite must adapt. Two strategies:

Pinned snapshots

Keep old product versions running in isolated environments. Agents train against v1.2, v1.3, v1.4 simultaneously. This catches regressions but requires infrastructure to run multiple product versions.

Adaptive evals

Update eval snapshots when the product changes. Mark old snapshots as deprecated. This keeps evals current but loses historical comparison.

Most teams use a hybrid: pin critical workflows, adapt the rest.

Ironclad cannot give OpenAI direct access to production. The sanitization pipeline must run inside Ironclad's VPC. The output (synthetic contracts and workflow snapshots) gets exported to OpenAI's training environment.

Key boundaries:

The sanitization pipeline is the trust boundary. If it leaks PII, the partnership fails.

Dimension Production Workflows Synthetic Benchmarks
Realism High (real complexity) Low (simplified tasks)
Data privacy Hard (requires sanitization) Easy (no real data)
Reproducibility Hard (version drift) Easy (static)
Failure diversity High (real edge cases) Low (designed cases)
Setup cost High (partnership required) Low (build in-house)

Production workflows expose real failure modes. Synthetic benchmarks are easier to control. The best eval suites use both.

Use this approach when:

Avoid this approach when:

The OpenAI-Ironclad partnership shows that production workflows can become agent training grounds. The plumbing is non-trivial: sanitization, versioning, failure labeling, and observability all need custom infrastructure. But the payoff is agents that handle real work, not just demos.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/openai-and-ironclad-…] indexed:0 read:5min 2026-10-07 Β· β€”