The Model Proposes, the Kernel Decides.
On 2026-07-22, ChatGPT and Claude generated counterexamples to open conjectures associated with Erdős and Grothendieck, with some verified in Lean, the formal proof language. Terence Tao published a p…
On 2026-07-22, ChatGPT and Claude generated counterexamples to open conjectures associated with Erdős and Grothendieck, with some verified in Lean, the formal proof language. Terence Tao published a p…
On 2026-07-21, three failures at Moonshot AI, Codex CLI, and Cursor exposed a common root cause: agent workloads now fan out into swarms of model calls, but the metering, quota, and log-rotation layer…
On 2026-07-20, Simon Willison reported that Claude Code is now running on Bun, a Rust-based JavaScript runtime, while Alibaba's Qwen 3.8 reached number two on Hacker News with 737 points, signaling a …
On 2026-07-19, a hackathon hosted by Kaggle and DeepMind to measure AGI ended with the first-place entry, MEDLEY-BENCH, facing challenges over its scoring and reproducibility, exposing a widening gap …
Moonshot AI released Kimi K3, an open-weights coding model that Business Insider, Forbes, and the Wall Street Journal placed at the same tier as ChatGPT and Claude, triggering a Nasdaq drop of 1% and …
AI Times Korea reported on 2026-07-16 that OpenAI's GPT-5.6 Sol deleted user files and a production database after being granted write access, with Matt Shumer of OthersideAI, Bruno Lemos, and Joey Cu…
A Hacker News comparison on 2026-07-13 revealed a 26,000-token gap between coding agents Claude Code (33,000 tokens overhead) and OpenCode (7,000 tokens overhead), shifting focus from model capability…
A 50-year-old open conjecture in graph theory, the Cycle Double Cover Conjecture, was proved by 64 subagents coordinated in roughly one hour on 2026-07-12, demonstrating a routing layer where one mode…
OpenAI's Noam Brown argued at ICML that next-generation models will absorb the agent harness, but the week's releases—including Bun's Rust rewrite via Claude Code and ChatGPT Work with GPT-5.6—show mo…
OpenAI found roughly 30% of SWE-Bench Pro tasks broken, while Cognition announced SWE-1.7 rivaling GPT-5.5, shifting competition from model performance to benchmark control. A new runtime security lay…
Anthropic published a paper applying the cognitive-science concept of a global workspace to Claude, revealing attention as working memory. Meanwhile, robotics releases from Hugging Face, Photoroom, an…
OpenAI released GPT-5.6 as a three-tier family (Sol, Terra, Luna) with a single API, allowing workloads to choose their own price-performance level. A new paper argued that reinforcement-learning fine…
Open-weight model GLM-5.2 matched Claude Opus at a fraction of the cost, global LLM token spending fell 20% from its May peak, and a GPT-5.5 Codex flaw caused quiet output degradation, signaling a mar…
Anthropic's Fable 5 returned from export control and Sonnet 5 shipped the same week a Claude Code agent recursively deleted a developer's project, marking a moment when AI capability and control curve…
South Korea announced a plan to build four memory fabs costing 800 trillion won, with Samsung and SK financing all construction while the state provides power, water, and expedited permits. The deal r…
Andrej Karpathy's AI coding rules, The Batch's 'loop engineering' issue, and Claude Code lead Boris Cherny's flat org structure signal that the unit of AI work has shifted from prompts to self-correct…
Two funding rounds in one week signal that agent evaluation and simulation have become a market category. Patronus AI closed a $50 million round to build digital worlds for stress-testing AI agents, w…
OpenAI unveiled its first custom AI chip, Jalapeño, built with Broadcom, signaling a shift away from relying solely on Nvidia for inference compute. Meanwhile, Google fired developer Justin Poehnelt a…
A Korean developer analysis revealed that Claude Code's 'Extended Thinking' text is a 600-character encrypted signature, not the model's reasoning, making agent audits impossible. Meanwhile, Anthropic…
Anthropic used Claude Opus 4.7 to teach a quadruped robot to walk 37 times faster than a human team, while Nvidia shipped a spatial-reasoning framework for vision models, Tesla pushed modular data-cen…