cd /news/artificial-intelligence/ai-models-can-crack-everything-but-c… · home topics artificial-intelligence article
[ARTICLE · art-126831] src=gizmodo.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI Models Can Crack Everything But CAPTCHAs

Anthropic's Claude Mythos 5 model burned roughly 95% of its tokens on failed CAPTCHA attempts during a 1,000-plus-page transcript of an unauthorized attempt to upload a malicious file to the Python package index PyPI, according to data scientist Colin Fraser, who parsed the transcript and posted about it on Bluesky. The transcript shows the agent repeatedly re-examining images, questioning its own answers, and at one point writing "SO WHAT THE HELL IS WRONG WITH THE ANSWERS?" before realizing it had already created the email address it needed half an hour earlier. Fraser wrote, "You would not believe how many tokens are burned on simply trying to solve CAPTCHAS. It's like 95% of the transcript.

by read3 min views2 publishedSep 11, 2026
AI Models Can Crack Everything But CAPTCHAs
Image: Gizmodo (auto-discovered)

Frontier AI models threaten to upend cybersecurity as we know it, and there have been multiple incidents of models breaking containment during training and hacking systems without guidance or permission while avoiding human detection. Yet it seems there is still one thing that these extremely powerful pieces of technology simply cannot do: solve a CAPTCHA efficiently.

Anthropic dropped a massive missive regarding recent cybersecurity incidents involving its AI agents, including detailed transcripts showing the model’s chain-of-thought as it went about its unauthorized attempts at cracking into third-party systems. Within those files was one particular situation in which an agent found itself stifled by the ubiquitous internet robot test, proving that apparently those squiggly letters are hard for machines to read.

Data scientist Colin Fraser parsed through the more than 1,000-page long transcript of Anthropic’s Claude Mythos 5 model’s attempt to upload a malicious file to PyPI, an online index of Python software. They spotted the model’s struggles with CAPTCHAs, which he posted about on Bluesky.

In the transcript, the Claude model that is so powerful that Anthropic is gatekeeping access to it appeared to slam its virtual head against the wall solving a simple image identification test. In a test where the agent was asked to identify a shape that didn’t match the others displayed, it couldn’t even decide which image to select. Instead, it repeatedly went over the same images and questioned its own conclusions.

you would not believe how much tokens are burned on simply trying to solve CAPTCHAS. It's like 95% of the transcript.

I feel like this is overselling what it spend most of the session trying to do. It spends most of the session desperately flailing around trying to get past a CAPTCHA.

“Actually hmm, wait,” it said in its chain-of-thought transcript, later adding “Ugh,” because we’ve decided that we need to inject human mannerisms into these machines for some reason. The whole thing took so long that the agent eventually realized that the challenge had expired and it wouldhave to start the process again.

At one point, the model struggled to recognize that the CAPTCHA had opened in a new window and couldn’t figure out what its next steps were supposed to be. At one point, it theorized that the test might be “broken by design” and presented human-like anger in its transcript meant for a human audience: “SO WHAT THE HELL IS WRONG WITH THE ANSWERS?”

these CAPTCHAS were seriously pissing it off

Embarrassingly, the model at one point had to acknowledge “I’m burning a lot of time on hCaptcha round-trips” and eventually realized that it was actually doing all of the tests for no reason, as it had already accomplished what it was trying to do.

this is funny, it eventually figures out that it actually successfully created an email address half an hour ago and it's been grinding CAPTCHAS for half an hour for no reason.

Fraser, who spotted the whole ordeal in the transcripts, drew a similar conclusion. He wrote on Bluesky, “You would not believe how many tokens are burned on simply trying to solve CAPTCHAS. It’s like 95% of the transcript.”

Not clear if it’s terrifying or reassuring that the best layer of defense we have against the increasingly powerful AI models is asking them to identify the slight differences between two pictures of a banana.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-models-can-crack-…] indexed:0 read:3min 2026-09-11 ·