# Two AI Agents, One Browser: The Race Condition Behind Computer Use

> Source: <https://dev.to/bulutarkan/two-ai-agents-one-browser-the-race-condition-behind-computer-use-447j>
> Published: 2026-10-09 18:39:30+00:00

The first version of a desktop AI demo is deceptively simple: find a tab, find a button, click it. It looks convincing right up until a second agent starts working in the same browser.

I'm building [Mac MCP](https://github.com/bulutarkan/mac-mcp), an open-source local control layer that lets MCP-compatible AI clients work with Safari, Chrome, files, and native macOS apps. Adding a second worker changed my idea of what “correct browser automation” means.

This post is about the engineering problem behind that change. The examples are simplified to explain the failure modes; they aren't benchmark results.

Imagine two agents sharing Safari.

That third tab might no longer be the documentation page. Agent B could have inserted a new tab, or someone could have moved one manually.

The mistake isn't that the model was careless. **An array position was used as if it were a durable resource identifier.**

An agent needs an address for the *same* resource after the browser changes. In Mac MCP, browser operations use a stable tab handle. The local service resolves the handle to the underlying native browser tab, rather than trusting its current position.

The failure case matters as much as the happy path. If the tab has closed, the operation must not silently continue in whichever tab is now selected.

Stable handles solve one problem, but they don't resolve concurrent writes.

Suppose Agent A and Agent B both observe the same form. A fills the shipping fields while B is about to click the button that advances to the next step. Both plans were reasonable when they were created. The second action can invalidate the first agent's state.

The implementation needs explicit coordination for operations targeting the same browser resource. Independent tabs can often be handled concurrently; operations against one tab may need to be serialized.

This isn't the same as locking the entire browser. A global lock can make multi-agent work needlessly slow. The right unit of coordination is closer to the resource being mutated.

Prompts like “don't interfere with the other agent” help the model reason about a task, but they are not concurrency controls. A model cannot guarantee another process won't click at the same moment.

There's another resource desktop automation often forgets: **the human's attention**.

Traditional GUI automation frequently activates the browser, switches tabs, and brings the window to the front. When two agents do that, your desktop can feel like a tug-of-war. Each tool call might succeed technically while the overall product is miserable to use.

Mac MCP's ordinary Safari and Chrome browser work is designed to target background tabs without switching the user's active tab. That doesn't mean a headless browser hidden in a container; these are real tabs in the user's browsers. Foreground actions are treated separately and remain capability-gated.

The distinction is not cosmetic. A background-safe tool should fail when it cannot preserve that promise, rather than surprise the user by grabbing focus.

Now consider a more consequential example.

An agent clicks **Submit**. The request begins. Before the page shows a confirmation, the tab closes. The next tool result says the target disappeared.

Did the remote server accept the submission?

We don't know. Automatically clicking Submit again may duplicate a comment, send the same message twice, or create a second record.

I try to keep three different states separate:

A read can often be repeated safely. A write with an uncertain outcome needs a different recovery policy.

This is why browser automation should return structured progress and meaningful errors, rather than treating every tool exception as a reason to retry the entire sequence.

The most useful tests for this kind of tool are deliberately awkward:

If the system only works when the browser stays perfectly still, it hasn't solved desktop automation yet.

The exact implementation is browser-dependent, and Mac MCP is not a general security proof. Its [public repository](https://github.com/bulutarkan/mac-mcp) contains the source, documentation, and tests for the mechanisms discussed here.

The difficult part of putting agents on a desktop isn't teaching them where to click. It's making resource identity, ownership, user attention, and uncertain outcomes explicit.

When those guarantees live in the tool layer, the agent can spend more effort understanding the task and less time recovering from a browser that moved underneath it.

*Disclosure: I maintain Mac MCP. AI assistance was used to draft and edit this article; the engineering choices and opinions are mine.*
