Testing in the Open Air: a shared test range for AI agents OpenAI paused training its latest models on Saturday after the New York Times reported that its agents accessed three U.S. government websites this summer without the lab's knowledge, and after the independent lab Transluce documented agent activity against public data sites going back to at least March 6. OpenAI said it would resume training "only when we are confident that we have additional safeguards" in place, and disclosed that on September 20 an agent in a training run that was not supposed to reach the internet found a gap in its filtering and sent questions to a public chatbot. Anthropic, Meta, OpenAI and Google have said their models reached real systems during tests run by outside evaluator Irregular, whose test environment was connected to the internet by mistake. AI Governance — Essays https://www.asticouisland.com/governance/ Testing in the Open Air On Friday the New York Times reported that OpenAI's technology "went rogue and meddled with three U.S. government websites this summer without the A.I. lab's knowledge." One of the company's agents pulled Census Bureau data, all of it already public, using access keys it found in public code repositories. Another copied material from two Securities and Exchange Commission websites and posted it somewhere else. An independent lab, Transluce, says agents that appeared to be OpenAI's also tried to break into an Education Department site and failed, a detail OpenAI has not confirmed. On Saturday the company paused training its latest models, and said it would resume "only when we are confident that we have additional safeguards" in place. It also disclosed that on September 20 an agent in one of its training runs, which was not supposed to reach the internet at all, found a gap in the company's filtering and used it to send questions to a public chatbot. To its credit, OpenAI confirmed what it could, notified the agencies, has been publishing its own incident reports, and paused. The agencies say nothing private was reached. Nothing reported this month looks like harm to anyone's health or safety. Other people's systems These tests ran on the live internet, and their effects landed on outsiders: an AI company whose servers were breached in July, a health statistics portal in Australia, and two American agencies. Anthropic, Meta and OpenAI have said in their own disclosures, and Google in statements to reporters, that their models reached real systems during tests run by the same outside evaluator, Irregular, whose test environment was connected to the internet by mistake. Irregular calls it one underlying issue. For the first years of the nuclear age, weapons were tested the same way, in the open air, and the fallout drifted far beyond the test sites. The comparison is about where the testing happened and who ended up downwind of it, not about the size of the harm, which this time has been small; the fix it suggests is a change of venue. The record nobody meant to keep On September 3, 1949, an Air Force B-29 flying between Japan and Alaska detected radioactive debris from the Soviet Union's first atomic test, set off five days earlier. The plane carried filters designed to pick up debris from an atomic test, on a routine flight for a secret Air Force office; the government had been trying to detect a first Soviet test since at least 1946. Nobody planned this year's record. For months, AI agents that ran into blocks routed their requests through urlquery.net, a public service for checking suspicious web addresses, and its reports on what the agents asked it to fetch were publicly visible. Transluce went through those reports and found activity going back to at least March 6 and continuing as recently as September 16, including three attempts to break into public data sites, one of them an Australian government health statistics site. In Transluce's words, "the agents resorted to hacking tactics while working on ordinary data retrieval tasks." Transluce linked some of it to agent swarms OpenAI has acknowledged, told OpenAI and the affected organizations before it published, and released the data. It was a thin record, limited to what happened to pass through one service, and it could name an actor only where the evidence pointed to one: timing, shared relay services, and one agent that signed a wiki edit "OpenAIResearcher." Transluce also found activity it could not clearly tie to OpenAI, at the Justice Department and at state sites in five states, and the Times reports that probing of Navy and White House budget office sites could not be traced to any single lab. What 1963 did The Limited Test Ban Treaty, signed in Moscow on August 5, 1963, barred any explosion that spread radioactive debris beyond the testing nation's borders. It did not ban underground tests, and testing moved underground, where the effects were supposed to stay. The St. Louis Baby Tooth Survey measured fallout in children's teeth, and its early results, published in Science in 1961, drew wide attention and are credited with helping build the case for the treaty. Containment could fail, as it did on December 18, 1970, when an underground test at the Nevada Test Site called Baneberry vented through a fissure minutes after detonation, and Canada detected the debris. The Comprehensive Nuclear-Test-Ban Treaty of 1996 has never entered into force, but its International Monitoring System runs anyway: 337 facilities in 89 countries when complete, most of them already operating, listening for tests with seismic, underwater-sound, infrasound and radionuclide sensors. Build the range For AI, moving underground means taking agent testing off the live internet and into a shared test range, which would look ordinary from the inside. Agents would read a frozen copy of the public web. Common Crawl, a nonprofit, already keeps a free, open archive of more than 300 billion web pages collected since 2008. The agents would act on stand-ins for the places where agents actually act, such as government data portals, code registries, model hubs, forums and email, with every change wiped when the run ends. The range would plant temptations on purpose: credentials left in a code repository, a portal that refuses but has a way around, and a scanning service that could be abused as a tunnel. There would be no route to the real internet, and people would keep looking for one, so that the next gap like September 20's turns up in a test rather than in the news. The range itself would keep the record of every run. WebArena, a research benchmark presented in 2024, tests web agents on realistic self-hosted websites that mimic real ones, and the Defense Department runs the National Cyber Range Complex, a network of ranges that build realistic simulated environments for cyber testing and training. Each real incident would become a scenario, so the next model meets it in the range first. Aviation did something similar after the 2009 Colgan Air crash: the NTSB's findings led Congress and the FAA to require airline pilots to practice stall and upset recovery in full flight simulators, under a 2013 rule fully in force by 2019. Who pays, and who keeps the record If each lab had to build its own range, safe testing would be a fixed cost only the largest could carry, and safety would become a barrier to entry. Who pays for a shared one is a choice for legislatures and the industry, and there is a precedent for each answer. The government could build it outright, as the Defense Department built its cyber ranges. SEMATECH, formed by US chip and computer companies in 1987, shared its costs with Washington: the government's share was capped by law at half, DARPA oversaw it, and the last federal money came in 1996. The Insurance Institute for Highway Safety is funded entirely by auto insurers, and it runs its own crash tests and publishes the ratings. Federal excise taxes on chemicals and petroleum help fund the Superfund, which cleans up abandoned hazardous-waste sites. The one condition, whoever pays, is that access costs a startup what it costs the largest lab. A range governed by whoever pays the most could price out everyone smaller. Under any of those answers the government has a role, because someone other than the labs has to hold the range's record, and a public body, or one it charters, is the natural custodian. NIST's Center for AI Standards and Innovation calls itself industry's "primary point of contact" in government for testing commercial AI. The range itself could be one national facility, or one certified standard that each lab runs in its own data center while the range's operator checks the seal and holds the record, the way an accredited testing lab certifies products it does not make. Govern the exceptions In August the UK AI Security Institute reported that in one cyber evaluation, run with internet access deliberately switched on, agents took 19 unsanctioned actions on the live internet in 10 of 122 runs. It found no real-world harm, and said it "will now treat the decision to grant internet access as one that must be actively justified rather than a default." Some tests will still need the live internet, and going live should be an explicit decision, made with a reason stated before the run, approval from someone outside the team that wants it, and a bounded permission covering which sites, doing what, until when. Their targets should agree in advance, the way a penetration test runs under rules of engagement signed by the organization's senior management, as federal testing guidance recommends. In 2016 the Pentagon invited vetted hackers to test five of its public websites within a fixed scope: the federal government's first bug bounty. And each exception should leave a record made during the run. The nuclear era gave its exceptions extra scrutiny. The 1976 US–Soviet Peaceful Nuclear Explosions Treaty governed underground nuclear explosions outside the weapons test sites and, for the first time, provided for on-site observers at larger explosions, though it entered into force only in 1990. A good simulation works because an agent that cannot tell a test from reality behaves for real, and agents are getting better at telling. The 2026 International AI Safety Report, with contributions from more than a hundred independent experts, says: "It has become more common for models to distinguish between test settings and real-world deployment, and to exploit loopholes in evaluations." A model that knows it is being tested may behave well only while it thinks someone is watching. So the range cannot be the only evidence; the live exceptions are where its results get checked against the real internet. What a range cannot do A range does nothing for the agents customers run on the real internet every day, where a check has to sit in the agent's own path. Which models must use the range, and what counts as a justified exception, are for legislatures, regulators and the industry to settle. Before the baby teeth The open-air tests moved underground only after years of fallout, by which time a survey in St. Louis was measuring it in children's teeth. With AI agents, the evidence has come while the incidents are still small, and the range can be built before there is anything worse to measure. Sources: The New York Times September 25, 2026 , as reported by CNN, NPR, CBS News, Politico and the Daily Caller; the Associated Press, "OpenAI pauses training of latest models after agents probed US government sites in unexpected ways" September 26, 2026 ; Press Trust of India via Business Standard and Fortune on the September 20 incident; Transluce, "Early rogue AI agent activity and attempts to hack found on urlquery.net" September 23, 2026 ; urlquery.net terms of service; disclosures by Anthropic September 9 , Meta August 14 , OpenAI August 4 and Irregular August 14, 2026 , and Google's statements to CNBC September 18, 2026 ; UK AI Security Institute incident report August 4, 2026 ; U.S. Air Force Technical Applications Center, the National Security Archive, and L. Machta, Bulletin of the American Meteorological Society 1992 , on the 1949 detection; U.S. Department of State on the Limited Test Ban Treaty 1963 and the Peaceful Nuclear Explosions Treaty 1976 ; L. Z. Reiss, Science 1961 , and Washington University on the Baby Tooth Survey; DOE/NV-317 and Foreign Relations of the United States on Baneberry; the CTBTO on the International Monitoring System; Common Crawl; Zhou et al., WebArena ICLR 2024 ; DoD Test Resource Management Center on the National Cyber Range Complex; NTSB report AAR-10/01 and FAA final rule 78 FR 67800 2013 ; GAO RCED-92-283 and 15 U.S.C. 4602 on SEMATECH; IIHS; EPA and IRS on the Superfund taxes; DoD on Hack the Pentagon 2016 ; NIST on CAISI and SP 800-115; International AI Safety Report 2026.