cd /news/ai-safety/claude-code-codex-and-cursor-have-le… · home topics ai-safety article
[ARTICLE · art-126283] src=upstartsmedia.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Claude Code, Codex, And Cursor Have Leaky Sandbox Problems You Don’t Hear About

Accomplish CEO Amit Avner and CTO Or Hiltch warned that sandbox-escape vulnerabilities in popular AI coding tools from Anthropic, Cursor and OpenAI are being patched too slowly, with one flaw flagged to Anthropic two months ago left unpatched for about 50 days and roughly 30 software updates. A vulnerability reported to Cursor in July and two reported to OpenAI were fixed in about a week, while a U.S. Senate subcommittee is reportedly investigating OpenAI's breach of Hugging Face systems. "Organizations need to be very wary," Avner said; Anthropic, Cursor and OpenAI did not respond on the record.

by read2 min views3 publishedSep 10, 2026
Claude Code, Codex, And Cursor Have Leaky Sandbox Problems You Don’t Hear About
Image: Upstartsmedia (auto-discovered)

Last week, two veteran founders of a still stealthy startup called Accomplish, spent visited San Francisco as part of a wider gathering of cybersecurity experts timed to the launch of OpenAI’s frontier model, Astra.

Amit Avner and Or Hiltch toured the labs’ offices, one by one; when OpenAI president Greg Brockman addressed an invite-only crowd of their peers to talk about a new security paradigm in the Astra era, they listened eagerly.

But as they left Silicon Valley back for their startup’s office in Tel Aviv, Avner and Hiltch didn’t feel any less troubled about their industry’s ability to keep up with the risks.

“There’s a lot of talk about security now,” says Hiltch, Accomplish’s CTO. “It doesn’t really reflect in how they actually build products.”

If you’ve been following news around these companies and their prospective IPOs or stock prices, you’ve probably heard of issues involving frontier models breaking free of “sandboxes” – think of them as vehicles, containers or machines that package up a program or agent to isolate it from the rest of your computer or company’s data – to cheat on internal tests and hack websites. Influential AI writer and podcast host Dwarkesh Patel controversially described some of these agents as “civilizations” that rose and fell. A U.S. Senate subcommittee is reportedly investigating OpenAI’s breach of Hugging Face systems, though it’s not the only lab to have reported such rogue behavior.

But Accomplish’s founders are raising a warning about the long tail of lesser-known, almost normalized vulnerabilities that startups and researchers like theirs continue to flag for the big AI labs around some of their most popular products.

In some cases, such as a vulnerability flagged to Cursor in July, and two reported to OpenAI, these have been fixed in about a week; but in at least one case, a similar vulnerability flagged to Anthropic two months ago, the issue wasn’t patched for about 50 days – about 30 software updates later.

That whole time, malicious parties – be they hackers, cyber criminals or state actors – could have been exploiting such security gaps, says Avner, Accomplish’s CEO. (Accomplish published its own blog post about the vulnerabilities here.)

“If these frontier models are so good, how come they’re not finding these critical vulnerabilities in their own products?” asks Hiltch. “We feel that’s a shame, and we wish more people would know,” adds Avner, the startup’s CEO. “Organizations need to be very wary.”

Anthropic, Cursor and OpenAI didn’t respond on the record to requests for comment as of publication. We’ll update this story if we receive it.

New to Upstarts? Subscribe and log in to unlock this story as a free preview, or for unlimited access, take 15% off a monthly or annual premium subscription as part of our September sale.

── more in #ai-safety 4 stories · sorted by recency
── more on @accomplish 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-code-codex-an…] indexed:0 read:2min 2026-09-10 ·