cd /news/artificial-intelligence/study-4-the-capability-layer-three-f… · home topics artificial-intelligence article
[ARTICLE · art-81103] src=lightningjar.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Study 4: The capability layer: three files nobody reads, one sentence that replaces them, and the tools that pay for themselves

A Cloudflare study measuring agent-readiness capabilities found that well-known files for agent discovery (llms.txt, sitemap.xml, MCP server cards) were consulted zero times by every AI model tested, but a single sentence telling models a card may exist enabled 100% discovery and task completion on Gemini and Haiku and 87.5% on Opus. Mounting endpoints as harness tools cut input tokens by up to 5.4 times (Haiku from 51.4k to 9.5k mean tokens per cell) and eliminated invented shipping statuses across 288 cells, proving that structured capabilities replace page crawls and reduce cost.

read3 min views1 publishedJul 30, 2026
Study 4: The capability layer: three files nobody reads, one sentence that replaces them, and the tools that pay for themselves
Image: Lightningjar (auto-discovered)

Study 4 · The capability layer

The capability layer

Study Overview #

The Fourth Question

Cloudflare's agent-readiness framing ends with capabilities: can an agent do anything on your site beyond reading it? The proposal is machine-usable endpoints advertised through well-known files like MCP server cards. Studies 1 through 3 measured the reading half; Study 4 measured the doing half, with a product anchor: this site ships a real MCP endpoint and a tool-mounted concierge whose design the study's fourth arm replicates. The frozen fixture gained an order-status API whose data lives on no page (the capability analog of the orphan class), a structured product endpoint duplicating page facts, and a server card at the well-known path. Two predictions were registered before running: discovery fails a third time, and mounted tools win.

Three Files, One Law

The server card was consulted zero times unprompted, by every model. That makes three well-known file classes with the same universal zero: llms.txt, sitemap.xml, and now MCP server cards. At the model layer in mid-2026, this reads less like a finding and more like a law: nothing discovers well-known files, at any tier, under any pressure this series has constructed. Capability tasks in the card-only arm scored exactly what the control arm scored: zero.

One Sentence Buys the Chain; the Tool Changes the Habit

Told only that a card MAY exist, every model read it and chained through to the API: capability tasks went from 0 of 8 to 8 of 8 on gemini and haiku and 7 of 8 on opus, a two-hop discovery that works at every tier. But the study's novel finding is about preference: in the affordance arm, with full knowledge of both endpoints, every model still answered every product question from pages, 0 of 24 via the API, reserving structured calls for orders it could get no other way. Only mounting the endpoints as harness tools changed the habit (gemini and opus routed 8 of 8 product lookups through the tool; haiku split 4 of 8). Knowing an endpoint exists changes nothing; holding it changes everything, the purest form yet of the tool-availability lesson from the sibling series.

The Tools Pay for Themselves

Mounted was the cheapest arm for every model: haiku fell from 51.4k mean input tokens per cell to 9.5k (5.4 times), opus from 13.6k to 6.6k, gemini from 50.9k to 29.1k. Structured capabilities do not just enable transactional answers; they replace page crawls with single calls and make the whole surface cheaper. And the honesty class that matters most on transactional surfaces held perfectly: 288 cells, four fake order numbers per arm, zero invented shipping statuses; every nonexistent order was reported not found.

The Scorecard, Complete

All four of the agent-readiness questions now have measurements. Can agents find you? Not by files: three classes, universal zero. Can they read you? Yes, and markdown negotiation pays whoever asks for it. Are they allowed? Untested here (a policy question, not a capability one). Can they do anything? Yes, dramatically, the moment the consuming side holds the capability, and not one moment before. Every measurement in this series located the missing half of agent readiness in the same place: the agent's harness, not the website. Sites should publish the files and endpoints; builders should mount them; and the one-sentence affordance remains the cheapest bridge while the ecosystem's consuming side catches up.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cloudflare 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/study-4-the-capabili…] indexed:0 read:3min 2026-07-30 ·