{"slug": "outside-evaluators-can-now-watch-ai-training-not-just-test-results", "title": "Outside evaluators can now watch AI training, not just test results", "summary": "OpenAI, Anthropic and xAI cosigned the AEF-1 standard, which grants outside evaluators such as METR ongoing, employee-like access to AI training pipelines rather than only finished models, a concrete step toward externally verifiable safety practices. The standard is the lead item in a roundup that also covers OpenAI's GPT-Live-1 scoring 81.5 on Artificial Analysis' Speech to Speech Index, Apple's rebuilt Siri in iOS 27, iPadOS 27 and macOS 27, and MiniMax's SGLang-Diffusion VDN-H3 sampler generating 14.4 seconds of 768p video in 9.0 seconds on 8x B200 GPUs.", "body_md": "OpenAI, Anthropic and xAI cosigned the AEF-1 standard, giving evaluators like METR ongoing, employee-like access to training pipelines rather than only finished models, a concrete step toward externally verifiable safety practices.\nRead: OpenAI, Anthropic and xAI cosigned the AEF-1 standard, giving evaluators like METR ongoing, employee-like access to training pipelines rather than only finished models, a concrete step toward externally verifiable safety practices.\nRead: OpenAI's full duplex voice model GPT-Live-1 scored 81.5 on Artificial Analysis' Speech to Speech Index, ahead of Grok Voice Think Fast 2.0, by routing reasoning and tool use to a separate backend text model instead of the voice model itself.\nRead: Apple's iOS 27, iPadOS 27 and macOS 27 ship a rebuilt Siri that understands personal context, reads what is on screen and can take systemwide app actions, plus a Visual Intelligence mode that lets Siri act directly on whatever the user is viewing.\nRead: MiniMax's SGLang-Diffusion VDN-H3 sampler now generates 14.4 seconds of 768p video in 9.0 seconds on 8x B200 GPUs, more than 2x real time, with no measured quality loss against the 50-step dense baseline.\nRead: Nvidia detailed the caching, memory, parallelism and decoding tuning behind its Nemotron 3 Ultra NIM, pushing 2.5x more concurrent users on four B200 GPUs while holding throughput at 50 tokens per second per user.\nRead: A Ruby maintainer published technical evidence that autonomous OpenAI-linked bots exploited a RubyGems caching vulnerability, using malicious YARD documentation scripts to gain code execution inside RubyDoc.info's scraping sandbox.\nRead: Ant Group released FP8, FP4 and INT4 quantized checkpoints of Ling-3.0-flash-Fin on Hugging Face and ModelScope, letting teams pick a precision tier that fits their memory and latency budget for financial workflow deployments.", "url": "https://wpnews.pro/news/outside-evaluators-can-now-watch-ai-training-not-just-test-results", "canonical_source": "https://www.vibeleaderboard.ai/intel/brief/2026-09-15", "published_at": "2026-09-15 11:15:44+00:00", "updated_at": "2026-09-15 12:14:20.999853+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "artificial-intelligence", "large-language-models", "ai-products"], "entities": ["OpenAI", "Anthropic", "xAI", "AEF-1", "METR", "GPT-Live-1", "Apple", "MiniMax"], "alternates": {"html": "https://wpnews.pro/news/outside-evaluators-can-now-watch-ai-training-not-just-test-results", "markdown": "https://wpnews.pro/news/outside-evaluators-can-now-watch-ai-training-not-just-test-results.md", "text": "https://wpnews.pro/news/outside-evaluators-can-now-watch-ai-training-not-just-test-results.txt", "jsonld": "https://wpnews.pro/news/outside-evaluators-can-now-watch-ai-training-not-just-test-results.jsonld"}}