Agent interactions are failing silently Merj R&D Director Giacomo Zecchini reported that in 100 purpose-built tests, browser agents failed silently on structural, naming, state and feedback defects that accessibility engineering already addresses — including one test where an agent clicked a misleadingly labeled "Cancel" button in 25 out of 25 runs when asked to save a document. Zecchini introduced Agent Interaction Monitoring, a task-level observability approach that starts a session on a defined task but lets the agent choose its own steps, and argued that public benchmarks and readiness scanners cannot tell a company whether agents can complete its critical journeys. The research, begun in November 2025 and previewed at Compound 2026 in May, concludes that accessibility and web standards are among the strongest existing foundations for agent-compatible interaction. Giacomo Zecchini https://merj.com/blog/author/giacomo-zecchini R&D Director · 44 min read Disclaimer We began this research in November 2025 and have watched agent behaviour change materially as models have evolved. A scenario that failed in December may pass today, and a new model may introduce a different failure. The observations below should therefore be read as version-specific evidence rather than permanent properties of AI agents. We shared a simplified early preview of this at Compound 2026 https://merj.com/events/compound-26 in May. Much of what follows argues that accessibility and web standards based implementation help AI agents. That is a side effect, not the main reason to do them. Accessibility exists so that people with a diverse range of abilities can use the web, and it should be done on that basis regardless of anything in this piece. Agents are a new actor on the web, but businesses have little visibility into what they do on their websites. In one of our tests, we deliberately gave two buttons misleading ARIA labels. When asked to save a document, one agent clicked Cancel first in 25 out of 25 runs. A site owner would be unlikely to spot that failure. Agents click, type, and attempt tasks in sessions that can resemble human traffic. If an agent takes the wrong action, there may be no ticket or reliable indication in analytics of what happened. So we built a way to watch, we call it Agent Interaction Monitoring. It delivers task-level observability for agents on real pages. Synthetic tests script every step and Real User Monitoring observes whatever sessions arrive. Agent Interaction Monitoring sits between the two: the session is one we start, on a task we define, but the steps are the agent's own. It records what the agent did, whether the intended state change actually occurred, and exactly where the sequence broke. We define the task and start the session, but the agent chooses its own steps. We used this approach in 100 purpose-built tests and then on live production pages, running multiple models repeatedly. For decades, websites have been designed for people and made discoverable to crawlers. Browser agents introduce a third kind of visitor: one that must also use the interface. We call this the interaction layer . Our findings point to two conclusions: 1. A large and commercially important subset of browser-agent failures originates in the same structural, naming, state and feedback defects that accessibility engineering already addresses. Accessibility is therefore one of the strongest existing foundations for agent-compatible interaction. 2. Public benchmarks and readiness scanners don't tell a company whether agents can complete its critical journeys. Only task-level measurement on your own site does. The problem you can't see Traditional web measurement often relied on a clear, though never perfect, distinction between human browser sessions and automated crawlers. Agents weaken that distinction because they might operate a full browser, run JavaScript, retain cookies, and produce user-like interaction sequences. Some identify themselves, some don't, and some are partially controlled by a person inside an existing browser session. Agentic browsing produces something that fits neither box - It's close to human without being human. No person is moving the mouse, but the mouse is moving. - It's unlike a conventional bot. Rather than crawling for an index, it's trying to complete a specific task. - It's hard to identify. It arrives in a real browser, uses an ordinary browser user-agent string, runs your JavaScript, and accepts your cookies or carries cookies from a previous human session. - It behaves like a user. It navigates, scrolls, clicks, types, and retries. It looks like a user session because functionally it is one. - It breaks attribution. Where did that session come from? A chat interface with no referrer? A browser a human was controlling until a few minutes ago? A datacenter IP in a country your user has never visited? - It can contaminate business analytics and RUM. If the agent executes client-side instrumentation, it might enter funnels, experiments, and even RUM tools' Web Performance data. The analytics effect is easy to underestimate. An agent that retries a form eleven times may look like one unusually persistent user. At sufficient volume, these unsegmented sessions could distort funnel analysis and experiment results. This research therefore concerns a population that remains difficult to observe in the wild. That limitation applies throughout our findings. When an agent fails, nobody files a ticket The interaction layer is easily neglected because feedback works differently for people and agents. When a person encounters a broken interface, you may eventually hear about it. The feedback loop is slow and incomplete, but it exists. People abandon a journey and appear in funnel drop-off, email support, or raise a ticket with a useful description: the Add to Cart button doesn't do anything on mobile Safari . An agent usually provides no such feedback. Depending on how it is built, it may retry before giving up without reporting the cause. The user eventually sees "I couldn't complete that" and finishes the task themselves. Worse, they may never learn what went wrong and simply conclude that your site doesn't work with AI. There is no ticket, and no log line recording that the agent misidentified the primary action. The failure leaves behind nothing that marks it as one. Not all agents act the same way Agents reach your functionality by one of two main routes. API-based access is system-to-system. The agent talks to a structured, documented interface, via protocols like MCP. There's no browser, no rendering and no clicking involved: it calls a function with typed parameters and gets structured data back. Its virtue is stability, speed, and precision. Browser-based access is human-ish-to-system. The agent opens a real browser or uses an existing one , loads your real page, and interacts using computer vision, DOM parsing, or both, mimicking clicks and keystrokes the way a person would. Its virtue is universal access and zero setup. If you've worked on the web for a while, this should feel extremely familiar. Structurally it's the scraping-versus-API question we've been arguing about since roughly 2005. Do you consume a site through its clean documented interface, or parse its HTML and hope the class names survive the next redesign? The two approaches have different trade-offs: browser-based access is fragile, while API-based access suffers from limited availability. Most sites don't have public APIs, and those that do often omit key features or data that are available through the user interface. That limitation explains much of the history of scraping: people used it because the API did not expose what they needed. Previously, the choice between an API and scraping was made by a developer concerned with reliability and maintainability. Given a suitable API at a reasonable cost, they generally chose it. Over the past fifteen years or so, those incentives helped produce a substantial API layer across the web. In the agentic web, the choice is made by the end user optimising for effort , and that changes the answer. MCP is promising for developers and advanced users. But every integration introduces adoption steps: finding the connector, understanding it, authenticating, granting scopes and deciding whether to trust it. And connectors only solve part of the problem. For any of this to work at scale, you also have to answer how people find the right service, whether the business still gets any credit or just disappears behind the assistant, how it wins that customer back when they never visited the site, plus identity, payments, and a way to see what happened. Connectors make a start on logins and permissions, and leave the rest for later. The browser needs less setup because the site is already there, and it already answers most of those questions. Discovery is search. Branding is the page. Identity is the login you already have. Payments are the checkout form that's already there. It breaks more often, but it inherits fifteen years of infrastructure instead of waiting for something new to be built and adopted. It also matches what people actually do. Type into a box and ask the following: No integration, no connector, no login flow, no decision about which competing commerce protocols the restaurant's booking system happens to support. The browser wins not because it's good, but because it's the only path that doesn't need everything else solved first. Developers may favour reliability while users favour low setup effort. As a result, structured protocols can grow without replacing UI-driven interaction. Meanwhile, several overlapping approaches to agent discovery, interaction and commerce have appeared in a short period. They don't all solve the same problem, but from a site owner's perspective they still create a moving implementation target. Commerce is the clearest example: Google's Universal Commerce Protocol UCP , announced in January 2026 and co-developed with Shopify, Etsy, Wayfair, Target and Walmart, gives agents a standard way to discover what a merchant supports and complete checkout without touching the page, while OpenAI and Stripe's Agentic Commerce Protocol ACP covers similar ground and AP2 handles payment authorisation. We're not here to predict which protocol or interaction model will win. The narrower point is that the browser remains the most widely available fallback because the interface already exists. Your UI is now an API Browser-based agents use websites much as people do. Unlike crawlers or API clients, they rely on visual reasoning, can mis-click and need to recover from errors, but often lack the context and intuition that help a person navigate a poor interface. As a result, your UI is now a de facto interface for software , even though nobody designed it as one, wrote a spec for it, or told you when it changed. Every ambiguity present on the website - for example an icon-only button, a div that behaves like a button, a layout shift or a focus-trapping modal - can become a failure in an undocumented interaction contract. This raises a new question: how do pages behave when agents act on them? That is different from asking whether content can be crawled bot access, robots.txt and firewalls or whether it renders SSR and hydration . The interaction layer comes afterwards: the agent has reached the page and it has fully rendered, but can the agent actually use it? How agents actually read a page To reason about interaction failures, you need a model of perception. Browser agents are commonly described as structure-first, vision-first, or hybrid. These are useful analytical categories rather than clean product boundaries: implementations vary, change over time and may use different modalities at different steps. Structure-first The agent consumes HTML, the DOM, semantic markup, and the accessibility tree, the browser's structured representation built from roles, accessible names, states and relationships. This is the same tree screen readers have consumed for decades. The failure mode is semantic. If your design system ships a button as a