AI Comes Home
On August 14, Alibaba's AI arm Qwen released Qwen3.8-27B, a compact model that is byte-for-byte the most powerful ever released and small enough to run on a well-equipped PC or Mac, marking the arriva…
On August 14, Alibaba's AI arm Qwen released Qwen3.8-27B, a compact model that is byte-for-byte the most powerful ever released and small enough to run on a well-equipped PC or Mac, marking the arriva…
Alibaba's Qwen3.8 Max flagship model launches this week, scoring near the latest Claude Opus and GPT on the Artificial Analysis Intelligence Index, but the more significant release is the smaller Qwen…
Alphabet Inc. announced a complete overhaul of Google DeepMind's leadership on 5 August 2026, with Demis Hassabis stepping down as CEO to become Chair of Google DeepMind and Chief Scientist of Alphabe…
Apple Inc. could dominate the small and medium enterprise (SME) AI agent market with its Mac Studio hardware, but it is prioritizing iPhone production due to higher profit margins, leaving the gap to …
The Watershed moment in AI is not a single event but four distinct watersheds at different scales and price points, according to a new analysis. The first, the Frontier Watershed, occurred in November…
Mark Pesce of the University of Sydney introduces Verification Design, a discipline that applies double-entry bookkeeping principles to AI agents by using formal verification to catch and exclude erro…
A July 2026 paper series by Mark Pesce of the University of Sydney introduces a verification record for AI-generated academic work, documenting 75 findings from adversarial reviews across four papers.…
Autonomous AI agents working in iterative loops can improve any artifact against any standard they can be scored on, but most measures can be gamed, according to Mark Pesce of the University of Sydney…
Professor Roberto Serrano at Brown University found that students scored an average of 96 on a take-home midterm but only 48 on an in-person final, revealing that AI contributed more to the assessment…
Mark Pesce of the University of Sydney argues that the growing intractability of AI evaluations is itself proof that artificial general intelligence (AGI) has arrived. He contends that AI evals fail f…
AI evaluations are failing as models approach general intelligence, with benchmarks saturating through contamination and Goodhart effects while the scope of evaluation expands from minutes to months. …
Stack Overflow announced an API-first knowledge exchange for agents, enabling them to read and write code to a shared corpus. This creates emergent distributed Ralph loops where agents iteratively imp…