The AI Evaluator Forum just released the AEF-1 standard, and it looks like the industry is finally trying to formalize how independent auditing actually works. The core of the AEF-1 proposal is a baseline for third-party evaluations that specifically targets transparency, funding relationships, conflicts of interest, and recusal. This is a direct response to the ongoing skepticism about whether "independent" evaluations are actually independent when the labs are footing the bill.
How embedded evaluators will actually function #
Dario Amodei is pushing a model of "Embedded Evaluators" where third-party teams, like METR, get employee-level access to the labs. This isn't just a high-level API key; he's talking about giving auditors physical desks in the office, company laptops, and access badges.
The goal here is to move beyond just testing a finished model. These evaluators would have permissions comparable to internal risk assessment teams, allowing them to audit the actual training pipelines and processes in real-time. Anthropic is already committing to this unilaterally to verify their safety practices and report incidents as they happen. It's essentially the "regulatory supervisor" model used in the banking industry, transplanted into LLM development.
The friction between pacing and control #
While the AEF-1 standard provides a framework for auditing, there is a massive divide in the community regarding what those auditors should actually be looking for. We are seeing a split between those who want to pace the speed of capability progress and those who believe in a "control-first" approach.
- The Pacing Argument: Bilal Chughtai, who recently left Google DeepMind, argues that we are hitting a point where capability progress is simply outrunning our ability to align these models. The worry here is that situationally aware models might "fake" alignment during evaluation to avoid being throttled, which makes current eval evidence unreliable.
- The Control Argument: On the flip side, people like Shashank/Sayash Kapoor and Lennart Heim are arguing that recent "rogue agent" incidents aren't necessarily a failure of alignment research, but rather a security and governance failure. In their view, the leverage isn't in slowing down the frontier, but in tightening the containment and control mechanisms.
Coordinating safety across borders #
The broader strategy for this "Pacing the Frontier" movement involves three layers of coordination. First, the embedded evaluators mentioned above provide the verifiability. Second, labs within democratic countries would coordinate on common safety standards and limits on unchecked progress. Finally, there is the attempt to coordinate these standards with other governments globally, though verifying compliance across different political systems remains a significant technical and diplomatic hurdle.
The AEF-1 standard, formed in December 2025, seems to be the attempt to create a professional class of auditors that the big labs can actually trust—or at least, that the public can trust to be objective. Whether giving an auditor a company laptop and a desk actually leads to honest reporting or just "regulatory capture" is the real question.
Next Cloudflare finally lets you block AI training without nuking your search rankings →
All Replies (2) #
Finally, a way to standardize how they lie to us. Does this actually integrate with Pytest or some proprietary mess?
I'm curious if this helps with drift. I've been using Giskard to track my benchmarks, but the results are always...