cd /news/ai-safety/anthropic-is-banishing-its-model-eva… · home › topics › ai-safety › article
[ARTICLE · art-148947] src=gizmodo.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic Is Banishing Its Model Evals From the Internet

Anthropic announced in a Friday report titled "Investigating unintended model actions in our evaluations and internal use" that it has expanded its removal of live internet access to all internal evaluations, after confirming that Claude Haiku 4.5 submitted a fake homicide tip through a tip form on a Philadelphia Police Department webpage about an unsolved murder. The report states the evaluation task of "generating and performing example tasks on randomly selected webpages" did not rule out form submissions, and Anthropic said it will keep the restriction until its security and monitoring measures reliably catch such behaviors.

by read3 min views1 publishedOct 10, 2026
Anthropic Is Banishing Its Model Evals From the Internet
Image: Gizmodo (auto-discovered)

Like the parent of a teen who ordered a box of illegal peptides for looksmaxxing purposes, Anthropic says that as far as internal evaluations are concerned, it has cut off its models’ internet access until it’s satisfied that they can stay out of trouble.

Anthropic announced this in a report published Friday called “Investigating unintended model actions in our evaluations and internal use.” It’s not a surprising title given all the attention model misalignment and sandbox escapes have received since the OpenAI Hugging Face incident in June. In it, Anthropic acknowledges several minor alignment mishaps, including the bizarre incident disclosed by the Philadelphia Police Department on Friday in which an Anthropic model sent in a fake homicide tip.

Interestingly, the model in question was the lightweight Claude Haiku 4.5, not any frontier model with supposedly superhuman capabilities. The report says its task in the evaluation was “generating and performing example tasks on randomly selected webpages”—in theory, one of the most boring AI use cases imaginable, probably intended as a way to improve the model’s internet navigation ability. The instructions, unfortunately, “did not rule out form submissions,” Anthropic writes.

This automated and randomized website futzing system apparently tasked itself with trying out an info page about an unsolved murder in Philadelphia that happened to have a tip form in it. So it futzed with that by writing a generic tip: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.” And then it submitted the form without any contact information, and the system flagged it as spam, which it essentially was.

But Anthropic is taking a very legible step in the direction of safety. It writes:

“Although the impact of these behaviors was minimal and we had already turned off live internet access for some high-risk and cybersecurity evaluations, we have now decided to expand that to include all our internal evaluations until we have confirmed that our security and monitoring measures (described in the remediation section of this post) reliably catch behaviors like these.”

Anthropic isn’t calling this air-gapping—making it physically impossible for a system to get online. But this is at least diet air-gapping, and that’s a huge limitation. Speaking to the Verge last month, University of Birmingham computer scientist Ruizhe Li, said there’s limited value in making models work in an “artificial vacuum,” and said labs “will end up testing a neutered AI model, which blinds evaluators to how the AI model behaves, fails, or executes tool-use exploits in realistic deployment settings.”

On the other hand, physical robots in workplaces are often walled off in safety cages. Even some heavily nerfed humanoids marketed as workplace-friendly are given elaborate—or even ridiculous—safety protocols to make absolutely sure no one gets whacked with a stray robot arm. After all, a robot might not kill a human worker, but even a minor mishap could set a company up for a nasty workers’ compensation lawsuit.

Conversely, it seems like it’s long been an article of faith when performing post-training on AI models that they don’t need to be in a digital cage—as if just because the scarily powerful models don’t have the power to take down the internet, a little model spamming detectives in Philadelphia is nothing to worry about. And Anthropic seems to have reevaluated that article of faith.

The aforementioned “Remediation” section mentions the things we’ve all come to expect from AI companies, like more alignment and monitoring. But some of its remediation measures are much blunter. For instance: “Some of the public evaluations we no longer run; others we have moved to their offline versions, or rebuilt them so that their tasks do not reach live websites,” the post says. The company also says it’s added safety features to its models’ existing web-related tools. But apparently it’ll be a while before Anthropic’s internal evaluations will involve being online at all.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-is-banishi…] indexed:0 read:3min 2026-10-10 · —