cd /news/ai-safety/outside-evaluators-can-now-watch-ai-… · home topics ai-safety article
[ARTICLE · art-130165] src=vibeleaderboard.ai ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Outside evaluators can now watch AI training, not just test results

OpenAI, Anthropic and xAI cosigned the AEF-1 standard, which grants outside evaluators such as METR ongoing, employee-like access to AI training pipelines rather than only finished models, a concrete step toward externally verifiable safety practices. The standard is the lead item in a roundup that also covers OpenAI's GPT-Live-1 scoring 81.5 on Artificial Analysis' Speech to Speech Index, Apple's rebuilt Siri in iOS 27, iPadOS 27 and macOS 27, and MiniMax's SGLang-Diffusion VDN-H3 sampler generating 14.4 seconds of 768p video in 9.0 seconds on 8x B200 GPUs.

read1 min views4 publishedSep 15, 2026

OpenAI, Anthropic and xAI cosigned the AEF-1 standard, giving evaluators like METR ongoing, employee-like access to training pipelines rather than only finished models, a concrete step toward externally verifiable safety practices. Read: OpenAI, Anthropic and xAI cosigned the AEF-1 standard, giving evaluators like METR ongoing, employee-like access to training pipelines rather than only finished models, a concrete step toward externally verifiable safety practices. Read: OpenAI's full duplex voice model GPT-Live-1 scored 81.5 on Artificial Analysis' Speech to Speech Index, ahead of Grok Voice Think Fast 2.0, by routing reasoning and tool use to a separate backend text model instead of the voice model itself. Read: Apple's iOS 27, iPadOS 27 and macOS 27 ship a rebuilt Siri that understands personal context, reads what is on screen and can take systemwide app actions, plus a Visual Intelligence mode that lets Siri act directly on whatever the user is viewing. Read: MiniMax's SGLang-Diffusion VDN-H3 sampler now generates 14.4 seconds of 768p video in 9.0 seconds on 8x B200 GPUs, more than 2x real time, with no measured quality loss against the 50-step dense baseline. Read: Nvidia detailed the caching, memory, parallelism and decoding tuning behind its Nemotron 3 Ultra NIM, pushing 2.5x more concurrent users on four B200 GPUs while holding throughput at 50 tokens per second per user. Read: A Ruby maintainer published technical evidence that autonomous OpenAI-linked bots exploited a RubyGems caching vulnerability, using malicious YARD documentation scripts to gain code execution inside RubyDoc.info's scraping sandbox. Read: Ant Group released FP8, FP4 and INT4 quantized checkpoints of Ling-3.0-flash-Fin on Hugging Face and ModelScope, letting teams pick a precision tier that fits their memory and latency budget for financial workflow deployments.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/outside-evaluators-c…] indexed:0 read:1min 2026-09-15 ·