The Pit Crew
Mark Pesce, a technology writer, reported that he used AI agents to optimize his local Qwen3.8-27B model, achieving a 20% speed improvement after GPT-5.6 Sol in Codex spent four hours tuning it. He al…
Mark Pesce, a technology writer, reported that he used AI agents to optimize his local Qwen3.8-27B model, achieving a 20% speed improvement after GPT-5.6 Sol in Codex spent four hours tuning it. He al…
A programmer spent about $260 in AI tokens to jailbreak an abandoned Amazon Fire tablet, and AI is being used to repurpose off-patent drugs for new diseases, illustrating a new 'cognition surplus' tha…
Meta unveiled Muse Glimmer, Alibaba released Qwen3.8-27B, and startup Ornith dropped Ornith-1.5-35B, a mixture-of-experts model that runs efficiently on consumer hardware, enabling agents to install a…
On August 14, Alibaba's AI arm Qwen released Qwen3.8-27B, a compact model that is byte-for-byte the most powerful ever released and small enough to run on a well-equipped PC or Mac, marking the arriva…
Alibaba's Qwen3.8 Max flagship model launches this week, scoring near the latest Claude Opus and GPT on the Artificial Analysis Intelligence Index, but the more significant release is the smaller Qwen…
Mark Pesce of the University of Sydney introduces Verification Design, a discipline that applies double-entry bookkeeping principles to AI agents by using formal verification to catch and exclude erro…
A July 2026 paper series by Mark Pesce of the University of Sydney introduces a verification record for AI-generated academic work, documenting 75 findings from adversarial reviews across four papers.…
Autonomous AI agents working in iterative loops can improve any artifact against any standard they can be scored on, but most measures can be gamed, according to Mark Pesce of the University of Sydney…
Mark Pesce of the University of Sydney argues that the growing intractability of AI evaluations is itself proof that artificial general intelligence (AGI) has arrived. He contends that AI evals fail f…
AI evaluations are failing as models approach general intelligence, with benchmarks saturating through contamination and Goodhart effects while the scope of evaluation expands from minutes to months. …
Software engineers at a recent conference expressed fear and grief over the spread of agentic AI systems, and discussed practical responses including token-cost management and cheaper diffusion models…