Claude Code, Codex, And Cursor Have Leaky Sandbox Problems You Don’t Hear About Accomplish CEO Amit Avner and CTO Or Hiltch warned that sandbox-escape vulnerabilities in popular AI coding tools from Anthropic, Cursor and OpenAI are being patched too slowly, with one flaw flagged to Anthropic two months ago left unpatched for about 50 days and roughly 30 software updates. A vulnerability reported to Cursor in July and two reported to OpenAI were fixed in about a week, while a U.S. Senate subcommittee is reportedly investigating OpenAI's breach of Hugging Face systems. "Organizations need to be very wary," Avner said; Anthropic, Cursor and OpenAI did not respond on the record. Last week, two veteran founders of a still stealthy startup called Accomplish, spent visited San Francisco as part of a wider gathering of cybersecurity experts timed to the launch https://www.reuters.com/legal/litigation/openai-launches-new-astra-model-amid-growing-scrutiny-over-agents-safety-2026-09-03/ of OpenAI’s frontier model, Astra. Amit Avner and Or Hiltch toured the labs’ offices, one by one; when OpenAI president Greg Brockman addressed an invite-only crowd of their peers to talk about a new security paradigm in the Astra era, they listened eagerly. But as they left Silicon Valley back for their startup’s office in Tel Aviv, Avner and Hiltch didn’t feel any less troubled about their industry’s ability to keep up with the risks. “There’s a lot of talk about security now,” says Hiltch, Accomplish’s CTO. “It doesn’t really reflect in how they actually build products.” If you’ve been following news around these companies and their prospective IPOs or stock prices, you’ve probably heard of issues involving frontier models breaking free of “sandboxes” – think of them as vehicles, containers or machines that package up a program or agent to isolate it from the rest of your computer or company’s data – to cheat on internal tests and hack websites https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/ . Influential AI writer and podcast host Dwarkesh Patel controversially described https://www.dwarkesh.com/p/openai-huggingface some of these agents as “civilizations” that rose and fell. A U.S. Senate subcommittee is reportedly https://www.axios.com/2026/09/10/openai-hugging-face-senate-investigation-hawley investigating OpenAI’s breach of Hugging Face systems, though it’s not the only lab to have reported such rogue behavior. But Accomplish’s founders are raising a warning about the long tail of lesser-known, almost normalized vulnerabilities that startups and researchers like theirs continue to flag for the big AI labs around some of their most popular products. In some cases, such as a vulnerability flagged to Cursor in July, and two reported to OpenAI, these have been fixed in about a week; but in at least one case, a similar vulnerability flagged to Anthropic two months ago, the issue wasn’t patched for about 50 days – about 30 software updates later. That whole time, malicious parties – be they hackers, cyber criminals or state actors – could have been exploiting such security gaps, says Avner, Accomplish’s CEO. Accomplish published its own blog post about the vulnerabilities here https://www.accomplish.ai/blog/beltdown-escaping-the-claude-code-sandbox/ . “If these frontier models are so good, how come they’re not finding these critical vulnerabilities in their own products?” asks Hiltch. “We feel that’s a shame, and we wish more people would know,” adds Avner, the startup’s CEO. “Organizations need to be very wary.” Anthropic, Cursor and OpenAI didn’t respond on the record to requests for comment as of publication. We’ll update this story if we receive it. New to Upstarts? Subscribe and log in https://www.upstartsmedia.com/subscribe to unlock this story as a free preview, or for unlimited access, take 15% off a monthly or annual premium subscription https://www.upstartsmedia.com/fallsale as part of our September sale.