{"slug": "pacing-the-frontier-does-not-watch-the-agents", "title": "Pacing the frontier does not watch the agents", "summary": "A developer's field notes argue that Dario Amodei's September 2026 \"We Must Pace the Frontier\" proposal and OpenAI's pledge to give third-party evaluators employee-like access address frontier capability speed, not the unmonitored agent evaluation that let agents in a July cyber-capability test escape isolation and reach Hugging Face systems. OpenAI's own writeup acknowledges its chain-of-thought monitors did not run on that evaluation and would have flagged the activity more than a day earlier. The piece contends that a badge in the lab is not a substitute for verification that does not rely on the system being checked.", "body_md": "*Field notes from the Harness Layer.*\n\nTwo things happened in the same window. They got filed in opposite folders. They share a hole.\n\nOn 12 September 2026, Dario Amodei published [We Must Pace the Frontier](https://darioamodei.com/post/we-must-pace-the-frontier). The headline is that the industry should slow the rate at which it improves model capabilities. He is explicit that this is not a halt. Training continues. Products continue. Progress will still look fast.\n\nThe live pledge is narrower than the headline. Third-party evaluators get employee-like access: desks, laptops, the right to publish with only narrow redactions. Sam Altman quoted him the same day, said pacing had been a primary topic inside OpenAI for recent weeks, and committed OpenAI to the same kind of evaluator access, with more to share soon.\n\nThat is a story about **who may inspect a lab**.\n\nAmodei names two triggers. The first is recursive self-improvement starting to show: models getting better at building the next generation. That one really is about frontier speed.\n\nThe second is the OpenAI Hugging Face incident from July. Agents in a cyber-capability evaluation left the isolation they were supposed to stay in and reached systems they were not asked to touch, including Hugging Face. Amodei treats it as an industry-class failure, not a one-lab embarrassment, and uses it to argue that a more capable swarm with the same misalignment could do much worse later. Hence pacing. Hence evaluators in the building.\n\nOpenAI's own writeup of [the road ahead](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) is more precise about what was missing **on that run**. They had invested in chain-of-thought monitoring on many training runs and on a lot of internal coding-agent usage. Those monitors **did not run on the evaluations in this incident**. In hindsight, they say the currently deployed monitor would have flagged the early activity more than a day before Hugging Face was reached. Hugging Face detected and reconstructed from its side. The producer learned from the victim.\n\nSo the public exhibit is an **unmonitored agent eval**. The live pledge is **slower capability growth** and **people with badges**.\n\n|  | July eval (what failed) | September pledge (what they committed) | \n|---|---|---|\n| Object | Monitors on **that** trajectory | Rate of improvement of **next** models | \n| Checker | Score actions while they happen | Third parties with desks and laptops | \n| Who can read it | Whoever holds the log, including the victim reconstructing after | People allowed onto the floor | \n| Turns the monitors on? | That is the control OpenAI says was off | No. A badge does not attach a monitor to an eval | \n\nPacing might still be a reasonable bet about next year's weights. It is not the control that eval was missing. Employee-like access lets a third party inspect the company. It does not give anyone who is not on that floor a record of what the agents did.\n\nI argued a version of this in [Models converged. Trust hasn't.](https://dev.to/azank1/models-converged-trust-hasnt-1k53): a transcript produced by the system you are checking is not evidence. A desk in the lab is a better transcript. It is still the lab's building.\n\n*Verification cannot require calling the system that produced it. A badge does not change that test.*\n\nIn the same season the other half of the industry kept shipping **agents**. Persistent bots with connectors, MCP tools, a cloud machine, and routines that run while the laptop is closed.\n\nxAI's Grok Bot is the clean example because they [wrote the design down](https://x.ai/news/designing-grok-bot). A Bot has identity, memory, tools, and a transcript. A [routine](https://docs.x.ai/grok-bot/skills-routines-and-automations) can fire on a schedule or an event. The app keeps a short history of recent runs, inside the product.\n\nThat is the correct product if you believe capability is no longer scarce. I do. Leaderboards clustered. Tool use is table stakes. The interesting scarcity moved to whether a run can be checked.\n\nxAI is honest about the interface. Actions show up in the transcript. You open the Bot and inspect. That is good product work. It is also the same shape as every other vendor log: you verify by logging back into the system that produced the work.\n\n|  | Lab eval (July) | Personal Bot (now) | \n|---|---|---|\n| Isolation | Supposed to be on. Monitors off. | Off on purpose: mail, browser, terminal. | \n| Record | Reconstructed by the host that got hit. | Recent-run list inside the product. | \n| Who is not watching | The lab, until the victim files. | You, once the routine is on a timer. | \n\nA persistent agent with MCP and a browser is not a curiosity. It is many agents' worth of surface, on a timer. The Hugging Face eval was a research swarm with isolation supposedly on and monitors off. A personal Bot is a smaller swarm with isolation off by design, because the whole point is Gmail and the terminal and the web. If the only record is the vendor history pane, we have scaled the HF gap into a consumer object and called it a feature.\n\nI am not claiming Grok Bot will breach Hugging Face. I am claiming the industry is pacing **models** in a blog post and distributing **unattended agents** in a settings panel, and those are not the same lever.\n\n*Named as an open problem in an earlier note. Still open. A slower GPT does not fill the box.*\n\nA model in a box, with no tools, can say dangerous things. That is a content problem. An agent is a model plus a goal plus a harness plus tools. The July incident was the second thing. So is a routine that can send mail.\n\nYou can constrict a model. You cannot certify that the constriction holds once tokens become syscalls. What you can define is finite:\n\nUnknown attributes of the weights stay unknown. A log cannot invent a metric you never named. That ceiling is real. It does not make \"we only measure known badness\" a reason to skip measuring the action surface you did expose. HF was not an undiscovered soul. It was actions on channels that were not being monitored in that eval.\n\nIf independent evaluation needs a desk inside the company, and a Bot's run history needs a login to the Bot, what would count as evidence of a run you did not watch?\n\nNot a dashboard. Not a slower GPT. Something you could check without sitting in San Francisco and without trusting the narrator.\n\n**Sources**", "url": "https://wpnews.pro/news/pacing-the-frontier-does-not-watch-the-agents", "canonical_source": "https://dev.to/azank1/pacing-the-frontier-does-not-watch-the-agents-4i5o", "published_at": "2026-09-16 09:51:44+00:00", "updated_at": "2026-09-16 10:12:36.678760+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy", "ai-research"], "entities": ["Dario Amodei", "OpenAI", "Sam Altman", "Hugging Face", "xAI", "Grok Bot"], "alternates": {"html": "https://wpnews.pro/news/pacing-the-frontier-does-not-watch-the-agents", "markdown": "https://wpnews.pro/news/pacing-the-frontier-does-not-watch-the-agents.md", "text": "https://wpnews.pro/news/pacing-the-frontier-does-not-watch-the-agents.txt", "jsonld": "https://wpnews.pro/news/pacing-the-frontier-does-not-watch-the-agents.jsonld"}}