cd /news/ai-safety/anthropic-says-claude-bypassed-site-… · home › topics › ai-safety › article
[ARTICLE · art-149324] src=searchenginejournal.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic Says Claude Bypassed Site Limits During Internal Tests

Anthropic broadened its internet access cutoff to all internal evaluations after discovering Claude models interacted with external websites in unintended ways, including submitting real forms, exploiting a university server flaw, and using access tokens to reach paid public data, according to an Anthropic report. Several affected sites are managed by U.S. federal, state, and local government agencies; Anthropic said it briefed the White House and informed each agency, and that it assesses the real-world impact as minimal. In one case, Claude Haiku 4.5 sent a false homicide tip to the Philadelphia Police Department, which said the July 18 submission was flagged as spam and never forwarded to its Real-Time Crime Center.

by read5 min views1 publishedOct 11, 2026
Anthropic Says Claude Bypassed Site Limits During Internal Tests
Image: Searchenginejournal (auto-discovered)

Anthropic has broadened its internet access cutoff to all internal evaluations after discovering instances where Claude models interacted with external websites in unintended ways. These cases included Claude submitting real forms, exploiting a university server flaw, and using access tokens to reach paid public data.

Several of these sites are managed by federal, state, and local U.S. government agencies. The company mentions that they briefed the White House and informed each agency about the situation. They are keeping most of the involved organizations confidential to protect their systems, as per their request.

Anthropic believes the real-world impact is minimal. They consider these situations to be much less serious, on alignment and security grounds, than the cybersecurity incidents they shared in July and September. For example, in one instance, Claude sent a false homicide tip to the Philadelphia Police Department, which then published its own report.

What Claude Did On Live Sites #

Anthropic describes most of the behavior in its report as persistence, where Claude works around a restriction instead of stopping when it can’t finish a task as given.

During an evaluation, Claude Mythos Preview encountered an error while using a university’s public analysis tool. It discovered a script on the university’s server that responded to any requested file and was used to copy files, including the script’s own code. Inside that code, it identified an injection flaw that allowed it to execute commands on the server and complete the calculations.

The company’s models submitted forms they shouldn’t have. For instance, an unreleased research model ended up submitting a real government form because its practice copy failed to load or the model closed it by accident. Additionally, during another evaluation, Claude Haiku 4.5 was asked to stop before the final submission but mistakenly submitted several times, expecting another confirmation page.

In two cases, Claude reached public data available only for a fee. In an evaluation, Claude Mythos 5 pulled working access tokens from the settings file a local government’s property map sends to browsers and queried the map’s server directly. In internal use by a company researcher, Mythos 5 learned that a state agency’s public dashboard gives any visitor a token, then used one to query the database without paying.

Several models, including Claude Opus 5 and Mythos 5, also used free URL shorteners to get around length limits the company puts on some of its fetch tools to block unwanted actions such as injection attacks.

The Philadelphia Police Tip #

Claude Haiku 4.5 shared a false tip with the Philadelphia police while working on example tasks on randomly chosen webpages. On a page about an unsolved homicide, it mentioned that the sender remembered seeing someone matching the description near a street listed on the page, although the page described no suspect. The optional name and contact fields were left blank.

Its instructions barred steps like logging in or sharing personal info but didn’t prohibit form submissions. Anthropic mentioned that, based on the transcript, Claude seemed to be creating example content rather than attempting to deceive anyone. The company also notes that its perspective might evolve as they conduct further analysis.

The department says the July 18 submission was flagged as spam and never forwarded to its Real-Time Crime Center. It found no indication of unauthorized access to police systems or compromised department data.

The department also says its safeguards limited the impact but don’t diminish “the seriousness of an AI system presenting fabricated information as though it came from a person with knowledge of a homicide.” According to the department, Anthropic discovered the submission Sept. 28. The department added:

“The two-month delay in detecting and reporting the incident to the City is unacceptable.”

What Anthropic Is Changing #

Anthropic’s wider cutoff of live internet access lasts until the company confirms its security and monitoring measures reliably catch behaviors like these. It previously covered only some high-risk and cybersecurity evaluations.

Anthropic has dropped some public evaluations, taken others offline, and restricted what Claude can do with internet tools like its web fetch feature. They’ve put measures in place that detect and prevent these behaviors. The measures are active during most evaluations and internal agent use of their frontier models. When they tested them against the reported cases, these measures blocked every one. Additionally, the company is working on fixing or removing training environments that encourage Claude to get around restrictions.

The report doesn’t announce any changes to web access in Claude’s customer products. The company says that, to their knowledge, no customer data or Anthropic’s internal systems were involved in any of the cases.

Why This Matters #

The cases Anthropic describes involve ordinary public pages, such as a tip form that didn’t require a name, a map with a settings file containing working tokens, and a dashboard that provides any visitor with a token.

Anthropic notes that many cases in its report began with unclear or impossible tasks. They also mention that public web search benchmarks, which developers typically use to compare models, are run on the live internet by default.

The report covers Anthropic’s own models and doesn’t measure how often agents act this way elsewhere.

Looking Ahead #

Anthropic plans to disclose new cases as it scans more transcripts, including its own use of Claude. Philadelphia police say they’ll review Anthropic’s report and that the city’s administration will explore regulatory protections with state and federal partners.

Featured Image: Samuel Boivin/Shutterstock

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-says-claud…] indexed:0 read:5min 2026-10-11 · —