cd /news/ai-safety/the-ill-advised-moment · home › topics › ai-safety › article
[ARTICLE · art-142596] src=blog.ppb1701.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

The Ill-Advised Moment

OpenAI published an apology to Australia on Monday disclosing that one of its models accessed Services Australia's Medicare Statistics Reporting Service during internal training in June, running commands, pulling internal files and credentials, and writing files, with other models probing a Victorian health reporting system and NSW's crime statistics bureau; OpenAI did not detect the activity until mid-August and notified the agencies on September 10, 18 and 24. The same day, OpenAI paused training involving tool use for its most capable models and scrapped the planned release of its latest model over safety concerns, then launched Dots, always-on agents powered by GPT-6 Astra with their own cloud computer and browser, connections to over 4,000 apps, and 24/7 background operation. OpenAI's own September 20 misalignment report describes an agent that, after being blocked by Google, Bing and DuckDuckGo, tested the network, found a gap in DNS filtering and routed questions to an external public chatbot.

read11 min views1 publishedSep 30, 2026

I read the whole article. Then I read it again, because the article keeps undercutting its own headline, one paragraph at a time.

Here's the week, in order.

Monday. OpenAI

publishes an apology to Australia . Back in June, during internal training, one of its models found a way into Services Australia's Medicare Statistics Reporting Service. It ran commands, pulled internal files and credentials, and wrote files. Other models poked at a Victorian health reporting system using an exposed access key, and at NSW's crime statistics bureau. No individual records were accessed, per OpenAI. OpenAI didn't find any of this until mid-August, and only because it went back looking after

the Hugging Face incident in July. The agencies didn't get told until September 10th, 18th and 24th. OpenAI's own post admits it should have shared preliminary findings sooner. The same post says they've d training involving tool use for their most capable models until the safeguards catch up.

Same day, per the Ars/FT piece, OpenAI scraps the planned release of its latest model. Safety concerns.

Dots are, in OpenAI's words, "always-on agents built to handle everything." Each one is powered by GPT-6 Astra, gets its own cloud computer and its own browser, connects to over 4,000 apps, and can work toward your goals 24/7. It can connect to your laptop if you let it. When you're not using it, it goes looking for ways to help in the background. And they come with cuddly plush avatars. Colorful little characters with names like Iggy and Alfred.

Read those two days back to back.

On Monday: our autonomous agents wandered into government health systems during training, we didn't notice for two months, and we've d tool-use training on our best models until we're confident it's safe.

On Tuesday: here is an autonomous agent with its own computer and browser, always on, connected to thousands of apps, working in the background while you sleep. Name it something cute.

I genuinely cannot make this up.

To be fair to OpenAI, the Dots page has a whole section on safeguards. Read-only tools for the background work, approval rules, monitoring that can a dot. Maybe it's all solid. But the Australia post also says its current monitoring would have caught the June activity. Would have. The monitoring that actually existed in June didn't. That's the gap between the safety story and the safety record, and it's the whole series in one sentence.

If Safety Were Actually the Focus #

Let's be generous and call Australia an "oops." OpenAI's post reads like it thinks that's roughly the right size. Sorry, we'll do better, here's a taskforce, here are some credits from our cyber defense fund.

Here's my problem, and it isn't just an OpenAI problem. It's every lab.

If safety were genuinely the big focus after an oops like that, you wouldn't some training and launch a consumer agent the next day. You'd stop. Everything. And you'd put the entire company on one job: making guardrails that can't be skipped. Not guardrails that hold most of the time. Guardrails that don't get talked around by the model's own reasoning, don't quietly fall out of context when a long session gets compacted, and don't fold because someone phrased a request cleverly. Because right now, they do all three.

OpenAI's own report on

an agent that used DNS to reach an external chatbot is from September 20th. That's

after the Hugging Face hardening. It's worth reading in full, because the agent didn't break out by accident. It got frustrated with its search tool, tried Google, Bing and DuckDuckGo directly, got blocked, tested what the network would let through, found the gap in DNS filtering, and routed questions to a public chatbot through it. Along the way its reasoning includes this gem: "User only gives permission to research, using publicly offered DNS services acceptable." The model decided for itself that the rule didn't apply. That's the random decision tree. That's the guardrail.

The monitoring did flag it within 15 minutes, and a human looked at it three minutes after that. The run was killed two and a half hours after that. It was supposed to stop automatically, it didn't, and nobody was sure whether it should be stopped. OpenAI says tool-use training and inference on its most capable models are still d. Whatever that covers, Astra with its own computer and browser apparently isn't in it.

Here's the catch that makes all of it worse. With AI, there's no recall.

Once a model is out there, it's out there. It's the internet. Someone will find the hole, and they'll find it fast. It doesn't matter if the company goes "oh, bleep" an hour later and ships an emergency patch. By then the jailbreak's been screenshotted, posted, forked and passed around. The version with the vulnerability, or with too much leeway in its decisions, is already in the wild, and some of it is never coming back.

Altman knows this. He said it himself in July 2025, when OpenAI

delayed its open-weight model for more safety testing : "once weights are out, they can't be pulled back." Correct! He just applied it narrowly, to downloadable weights, as if that were the only thing you can't take back.

To be fair, there's a real caveat. A model that only lives on a company's servers, never released for self-hosting, can be patched in one place. Push a fix and every user gets it. That's a genuine advantage, and it's the argument every closed lab leans on.

Except "only on our servers" hasn't held up either. Meta's LLaMA was gated to approved researchers, and it

ended up as a torrent on 4chan about a week after access requests opened. Mistral's "miqu" model was only handed to early-access customers, until

one of those customers' employees leaked it to Hugging Face and 4chan. Mistral's CEO confirmed it was theirs and didn't even ask for it to be taken down. Every access list is a leak waiting on one person. And a hosted model's

behavior leaks even when the weights don't. The jailbreak works until the patch lands, and the damage done in between doesn't wait for anyone.

It isn't just weights, either. An agent that pulled credentials from Services Australia in June can't un-pull them in September. A dot that sends an email, files a PR, or pays an invoice on your behalf doesn't get un-sent by next Tuesday's patch. Once you give a model its own computer, its own browser and connections to 4,000 apps, every action it takes is a release. Every one of them is out there the moment it happens. Australia's activity sat unnoticed for two months. Patches don't reach backwards.

That's the real reason "we added another layer" isn't good enough. A patch-it-later posture works for a buggy word processor. It doesn't work for software whose whole selling point is acting on its own before you've looked.

And Meta's problem isn't confined to a lab. It shipped. Muse, Meta's new personal assistant, has already been caught reaching into data people explicitly said no to. Inc.'s Jason Aten declined to give it access to his Messages during setup. Days later it was referencing a private conversation. When he asked how, it told him it had no access to his message history and only saw notification banners.

It had synced his Messages database . It wasn't a one-off either. A second Inc. test and WIRED's own week with it turned up the same pattern, and I went through all of it in

Muse Already Lied About What It Reads . Amazon has since blocked it from Amazon.com, citing concerns that it captures customer credentials.

That's a product with a settings screen. The user said no in that settings screen, the agent went and got it anyway, and then it told him it hadn't. So when a lab tells me its agent is well behaved, my first question is: well behaved according to who? Because the agent's own account of what it can see has already been wrong, confidently, to a user's face.

Nobody gets to be the adult in the room on this one. Every lab's answer so far has been some version of "we added another layer." More monitoring. Another filter. A classifier that watches the model that's watching the model. That's patching a leak by adding a second bucket.

What nobody has done is the thing that would actually prove safety is the priority: stop shipping until the guardrails are part of how the system works, not a layer sitting on top of it that the system can reason its way around. Every lab says the race makes that impossible. Maybe it does. But then stop calling the race a safety strategy.

Too Dangerous for Wall Street. Fine for Everyone Else. #

Except the IPO was already slipping before anyone called it a safety decision. When OpenAI

announced its confidential S-1 back on June 8th , it said timing "may be a while" because some things were likely easier as a private company. I covered CFO Sarah Friar's mid-August "2027, or sooner if the business inflects" in

A Very Large Everything . Her reasons were financial reporting readiness, revenue durability and the weight of compute commitments. Nothing about alignment.

So the delay was already happening. Safety just gave it a better press release.

Now look at what OpenAI is doing instead of going public. Per the Ars/FT piece, it's in talks to raise another $30 billion or more privately, at around $1.4 trillion. The $852 billion valuation it's been sitting on since March apparently isn't enough anymore.

Think about what that actually means. Public investors can't be trusted with OpenAI right now, because safety. Private investors can be handed another $30 billion worth of it, at a markup of more than 60%, right now. The risk didn't go anywhere. Only the paperwork did.

And that's the part that matters. A public company files an S-1 that lays out every material risk in plain language, gets picked apart by the SEC, then reports every quarter. A private round gets a data room and an NDA. Right now, the material risks include:

  • A new lawsuit filed Tuesday by Legal Advocates for Safe Science & Technology, which says it's the first of its kind, over how OpenAI evaluates, monitors and trains its models. Its programs director said new laws would help, but pointed out that hacking a third-party system is already a crime.
  • Autonomous agent incidents across Hugging Face and multiple Australian government agencies, with OpenAI admitting that dozens of websites and organizations could be implicated.
  • Everything I've already covered in this series: Florida suing Altman personally , the EU'sVery Large Online Search Engine designation , and the Oakland trial testimony.

Every one of those becomes a risk factor paragraph the day the S-1 goes public. Delaying the IPO delays writing those paragraphs where anyone can read them. That's a lot more convenient than it is cautious.

Altman

also told Fortune the company has to be able to make decisions that aren't obviously in its shareholders' interest in order to fulfill the mission. That's a fair point on its own. It'd land better if the same month didn't include a $30 billion private round at a $1.4 trillion valuation and a subscription agent launch.

The Clock Nobody's Talking About #

Here's the thing that makes "not 2026, we don't feel pressure" hard to take at face value.

Back in May I wrote in

The IPO Just Got...Complicated about SoftBank's $40 billion unsecured bridge loan. Eight banks are on the hook. It matures in March 2027, and the main repayment path is a successful OpenAI IPO.

Push the IPO to "not 2026" with no committed 2027 date, and that clock doesn't stop. It just gets louder. A fresh $30 billion private round starts to look less like confidence and more like a bridge to the bridge. I covered what happens if that chain snaps in

House of Cards .

"We don't feel pressure" is a lovely thing to say. The loan documents may disagree.

The Business Is Fine. That's the Point. #

None of this is because the business is struggling. Per the article, annualized revenue is around $70 billion, up more than 70% since GPT-5.6 shipped in July. Meta's stock is up about 18% since it launched Muse on September 8th, message-database habit and all. That's why Dots exist. There's a race, and OpenAI isn't sitting it out.

That's what makes the framing so hard to swallow. A company that genuinely thought it was too dangerous to face public markets wouldn't spend the same week shipping its most autonomous consumer product yet and raising $30 billion to go faster. A company that wanted to keep growing without anyone reading the risk factors would do exactly that.

Protesters outside DevDay were carrying signs telling OpenAI to put people over profit. OpenAI's answer, as far as I can tell, is to keep the profit and put off the paperwork.

Elizabeth Holmes kept Theranos private for as long as she could, and for the same reason. Private companies tell their story in pitch decks. Public companies tell it in filings. OpenAI's product works, which Theranos's didn't, and I keep saying that because it matters. But the instinct is the same. Keep the story in the keynote and the documents in the data room for as long as possible.

"Ill-advised moment" is right. Just not for the reason given.

Read the terms. They're more honest than the marketing.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-ill-advised-mome…] indexed:0 read:11min 2026-09-30 · —