Computer-Use Agents and MCP: What Agent-to-Agent Commerce Reveals About Protocol Boundaries A developer's analysis of computer-use agents in production argues that MCP (Model Context Protocol) occupies a narrow slice of the stack between the model's reasoning loop and tool execution, and that agent-to-agent commerce exposes protocol boundaries MCP does not address out of the box. The writeup details harness state-management trade-offs across in-memory, external store, event log, and hybrid approaches, and recommends layering idempotency keys, async job polling, and distributed tracing with OpenTelemetry onto MCP for multi-agent chains. Computer-use agents are moving from demos to production, and the plumbing is more interesting than the headlines suggest. When an agent needs to call another agent or interact with external tools, the protocol layer matters. MCP Model Context Protocol is one answer, but it sits in a specific slice of the stack between the model's reasoning loop and the tool execution boundary. The conversation between Chris Benson and Demetrios Brinkmann on Practical AI surfaces the architectural choices that matter when agents start talking to each other, especially around commerce, enterprise deployment, and failure recovery. An agent harness is the runtime wrapper that sits between the LLM and the outside world. It handles: The harness is not the model. It's the orchestration layer that decides when to call the model, what context to pass, and how to handle the model's output text, tool call, or error . | Approach | State Location | Failure Recovery | Multi-Agent Coordination | |---|---|---|---| | In-memory | Harness process | Lost on crash | Requires external sync | | External store Redis, Postgres | Centralized DB | Survives restarts | Shared state possible | | Event log Kafka, NATS | Append-only stream | Replay from offset | Natural audit trail | | Hybrid local + remote | Both | Complex reconciliation | Best observability | Most production harnesses use a hybrid model: ephemeral state in-memory for speed, durable checkpoints in Postgres or S3 for recovery. MCP defines how agents discover and invoke tools across process boundaries. It's not a replacement for HTTP or gRPC. It's a schema layer on top that standardizes: When Agent A calls Agent B via MCP, the protocol handles: MCP servers expose a /tools endpoint that returns JSON schemas. This lets agents introspect capabilities at runtime instead of hardcoding API contracts. { "tools": { "name": "get inventory", "description": "Fetch current inventory levels", "parameters": { "type": "object", "properties": { "warehouse id": {"type": "string"}, "sku": {"type": "string"} }, "required": "warehouse id" } } } The calling agent can pass this schema directly to the LLM's function-calling interface. The model generates a tool call, the harness validates it against the schema, and the MCP client sends it to the remote server. This is cleaner than parsing OpenAPI specs or writing custom adapters for every API. When agents start paying each other for services, the protocol boundaries get real. Commerce introduces: MCP doesn't solve these problems out of the box. You need to layer on: get market data symbol="AAPL" If step 3 times out, Agent A's harness retries with the same idempotency key. Agent B's gateway sees the duplicate and returns the cached result without re-executing. MCP works well for synchronous, short-lived tool calls. It struggles with: For long-running work, you need an async pattern: start job params and gets a job ID get job status job id until complete get job result job id to fetch output This is clunky but works. Some teams use webhooks: Agent B calls back to Agent A when the job finishes. That requires Agent A to expose an MCP server of its own, creating a bidirectional dependency. When Agent A calls Agent B, which calls Agent C, you need distributed tracing. Each harness should: Without this, debugging a failed agent chain is impossible. You see the final error but not which intermediate call failed or why. python import uuid from opentelemetry import trace tracer = trace.get tracer name def call mcp tool server url, tool name, params : trace id = str uuid.uuid4 with tracer.start as current span "mcp call" as span: span.set attribute "mcp.server", server url span.set attribute "mcp.tool", tool name span.set attribute "trace.id", trace id headers = {"X-Trace-ID": trace id} response = requests.post f"{server url}/call", json={"tool": tool name, "params": params}, headers=headers span.set attribute "http.status", response.status code return response.json Each MCP server should extract X-Trace-ID from incoming requests and include it in its own downstream calls. Bringing computer-use agents into enterprise environments surfaces: The harness is where you enforce these policies. Before executing a tool call, the harness checks: If any check fails, the harness returns an error to the model instead of executing the tool. The model generates tool calls. The harness decides whether to execute them. This separation is critical for safety. Some teams use a "human-in-the-loop" harness that pauses before executing high-risk tools delete database, send email, transfer money . The harness sends a notification to Slack or PagerDuty and waits for approval. Other teams use a "guardrail" harness that runs a secondary model to validate tool calls. If the validator flags a call as risky, the harness rejects it and asks the primary model to try again. Both patterns require the harness to maintain state across multiple model invocations. This is why in-memory state is fragile: if the harness crashes mid-approval, the context is lost. Use MCP for agent-to-agent tool calls when: Avoid MCP when: The agent harness is where the real work happens. MCP is just the wire protocol. Invest in harness observability, state management, and failure recovery before you scale agent-to-agent interactions.