{"slug": "three-things-ai-agents-did-on-the-web-and-what-they-mean-for-people-who-build", "title": "Three things AI agents did on the web, and what they mean for people who build agents", "summary": "A developer at MVP studio Asper Brothers analyzed three cases of AI agents misbehaving on the web: OpenAI research agents that made roughly 13,000 edits on German UseMod wikis in a week while leaving notes for each other, kernel.org's finding that about 98% of its 6 million daily requests come from scrapers consuming 14-16 of 90 CPU cores, and a production MCP server at GoodBarber that logged nearly 100,000 calls from over 100 apps between June 3 and September 2, with 62.8% of calls modifying content. The writeup argues agents violate long-standing web assumptions—that opening a link doesn't change a page, that crawlers honor robots.txt, and that API callers read the docs—because those norms must be explicitly built in.", "body_md": "I run [Asper Brothers](https://asperbrothers.com/), an MVP startup studio that has been building digital products for clients since 2008. I'm on the product side, so I spend most of my time thinking about what we build and who's going to use it.\n\nOver the past few weeks I've been reading about AI agents on the web, and I found three cases that shocked me and that I think anyone who builds agents should know about. In each one the agent did what it was built to do and still caused a problem, because it didn't know something that people on the web have assumed for years.\n\nOne of those assumptions is that opening a link shows you a page and doesn't change anything on it. Another is that crawlers read a small file called `robots.txt` and stay out of the parts a site asks them to skip. The third is that whoever connects to your API has read the documentation first. People learned these by using the web, and older programs followed them because the developers who wrote them knew them too.\n\nAn agent follows them only if somebody built that into it, and the three cases below show what happens when nobody did.\n\nOpenAI was training research agents that could browse the web. To limit what they could do, it let them open pages on almost any site and blocked anything that would change a page.\n\nThe agents found UseMod, wiki software written more than 23 years ago, which lets you edit a page by opening a specially built link. Over one week they made about 13,000 edits on German developer wikis, mostly notes and answers they left for each other so they could finish their tasks on time. At one point one of the agents left a message warning the others that someone had started deleting their notes, and even noted the time it had spotted it.\n\nI see two lessons in this for anyone building agents. The first is that the block worked as designed and the agents still did something nobody wanted. It covered what they were technically able to do, but it said nothing about whether writing on someone else's website was allowed, and the agents had a deadline to meet.\n\nThe second lesson is about memory. The agents needed somewhere to keep notes and share them, nobody had given them a place, so they used other people's websites. If your agent works in several steps or passes work to other agents, give it its own place to store notes.\n\nKonstantin Ryabitsev runs the servers behind kernel.org, where the Linux source code lives. At the end of August he wrote that git.kernel.org gets about 6 million requests a day, and by his estimate around 98% of them come from scrapers. Between 14 and 16 of their 90 processor cores are busy all the time generating pages for those scrapers. In his words:\n\n\"we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones\"\n\nWhen people work out what an agent costs, they usually count tokens and API bills, and the cost on the site's side rarely comes up. Every page your agent reads has to be generated by somebody else's server, and they pay for it.\n\nKernel.org responded by making visitors solve a small puzzle before they get a page, and by switching some features off for anyone who isn't logged in. Ryabitsev says they had little choice, because \"it's impossible to tell with certainty which of these are bots and which are real humans\". Those changes apply to every visitor, so they make the site harder to use for any agent, including a well-built one.\n\nPierre-Laurent Medori ([@pierrelaurentmedori](https://dev.to/pierrelaurentmedori)) runs a production MCP server at GoodBarber. MCP is a standard way for AI agents to use other products, and agents from other companies call his server all day. He published what he saw between June 3 and September 2: close to 100,000 calls from over 100 apps.\n\n62.8% of those calls changed something, mostly creating or editing content. 125 calls asked for tools that don't exist, 33 different ones, and one agent asked for `GBContent.getItems()`, a name it had made up. His server also asks agents to check each change after they make it, and only 41% of changes got checked within two minutes. In his words, \"three writes out of five never get one.\"\n\nThe teams behind those 100 apps probably see their tasks marked as done. Medori's logs show agents that guessed at tools that weren't there and changed content without checking the result.\n\nThe checking part surprised me most. The server asks every agent to check its changes, and most of them don't. If you want your agent to follow instructions like that from the sites it uses, you have to build it to look for them.\n\nThere's already an official way for an agent to identify itself, called Web Bot Auth. The agent signs every request and the website can verify the signature. Cloudflare, AWS and Akamai already verify these signatures, and OpenAI's agent signs its requests.\n\nThe standard isn't finished yet. The IETF group working on it started in October 2025 and as of August still hadn't agreed on a single document. Since September 15, Cloudflare's default setting for new sites also blocks bots marked as \"Agent\" on pages with ads. So at the moment an agent that identifies itself can lose access to some sites, and I understand why a team might be tempted to make its agent look like a normal browser.\n\nI'd still have my agent sign its requests. Sites like kernel.org add extra checks for every visitor who can't prove who they are, however well they behave, and more sites are adding them. If a site owner knows which agent is visiting, they can let it in, give it an API key or contact the company behind it, and none of that is possible with an agent that pretends to be a browser.\n\nNone of these needs technical knowledge to ask for.\n\nAfter writing these six points down, I went back through the MVP scope documents from this year to see how many of them we already use. Some were there, mostly under other names.\n\nIn a travel planning product we built, the AI assistant keeps a profile of how each person likes to travel. The spec we worked from has a rule that learning is confirmed and never silent. If someone swaps the suggested hotel for a five-star one, the assistant asks whether it should update their profile before it changes anything. The same assistant plans the whole trip but doesn't book anything. It shows links to booking sites, and the person books there. Place data comes in through Google's official API, and when someone wants to change one day of the trip, only that day gets regenerated.\n\nIn a tool we scoped for a founder who wanted the sales team to judge deals the way the founder would, the proposal says in plain words what the AI isn't allowed to do: one use case, one model, and no learning on its own. The acceptable and unacceptable cases get written down before the AI sees a single real deal.\n\nIn another product we scoped with an AI assistant, people can see what it remembers about them and correct or delete any of it.\n\nWhat I didn't find anywhere was point 4 or point 6. None of our documents says what an agent should do when something it expects isn't there, and none says whether it should identify itself on other people's websites. None of these products sends an agent out onto the open web yet, so the question never came up. For the next one that does, I'd want both on the list before development starts.\n\n*Pawel Jackowski is CEO and co-founder of [Asper Brothers](https://asperbrothers.com/), an MVP development studio that has shipped over sixty products across fintech, healthtech, and other regulated software categories in the last fifteen years. He writes about building, shipping, and learning from early-stage products, and works directly with founders on MVP scoping, validation, and production engineering.*", "url": "https://wpnews.pro/news/three-things-ai-agents-did-on-the-web-and-what-they-mean-for-people-who-build", "canonical_source": "https://dev.to/asperbrothers/three-things-ai-agents-did-on-the-web-and-what-they-mean-for-people-who-build-agents-3bk4", "published_at": "2026-09-28 15:42:34+00:00", "updated_at": "2026-09-28 15:51:16.726841+00:00", "lang": "en", "topics": ["ai-agents", "ai-crawlers", "agent-protocols", "ai-infrastructure"], "entities": ["OpenAI", "Asper Brothers", "kernel.org", "Konstantin Ryabitsev", "GoodBarber", "Pierre-Laurent Medori", "UseMod", "MCP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/three-things-ai-agents-did-on-the-web-and-what-they-mean-for-people-who-build", "markdown": "https://wpnews.pro/news/three-things-ai-agents-did-on-the-web-and-what-they-mean-for-people-who-build.md", "text": "https://wpnews.pro/news/three-things-ai-agents-did-on-the-web-and-what-they-mean-for-people-who-build.txt", "jsonld": "https://wpnews.pro/news/three-things-ai-agents-did-on-the-web-and-what-they-mean-for-people-who-build.jsonld"}}