Benchmark headlines break down: cost per task, not per token, decides the day Artificial Analysis reported that Qwen3.8 Max's per-token price fell while its cost per Intelligence Index task more than doubled versus Qwen3.7 Max, and its GDPval-AA Elo lead over Kimi K3 came from taking roughly four times more turns per task, highlighting that per-token pricing misleads. The same model regressed ten points on AA-Omniscience (+14 to +4) with hallucination rate roughly doubling at flat accuracy, and Artificial Analysis patched the Intelligence Index to v4.1.1 with changed graders, making cross-version comparisons invalid. Meta's Muse Spark 1.2 landed on the cost-per-task Pareto frontier, and Kimi K3 reached GA in Copilot before a rollout pause, while Agent Plugins shipped as a cross-client standard from OpenAI and Cursor. The day's most useful number was not a leaderboard rank but a denominator. Artificial Analysis showed Qwen3.8 Max cutting per-token price while more than doubling cost per Intelligence Index task, winning GDPval-AA Elo largely by taking roughly four times as many turns, and regressing ten points on AA-Omniscience as it stopped abstaining — three different ways a headline score can hide latency, spend and hallucination risk, arriving alongside an Intelligence Index grader change that makes cross-version comparisons invalid. Against that, Meta's Muse Spark 1.2 landed on the cost-per-task frontier and Kimi K3 reached GA in Copilot before a rollout pause, giving builders cheaper defaults whose real economics still need per-workload measurement. Underneath the model news, the plumbing converged: Agent Plugins shipped as a cross-client standard from OpenAI and Cursor on the same day the MCP ecosystem moved toward stateless servers and web-side tool interfaces. Debate: Qwen3.8 Max is the cleanest case yet that per-token pricing misleads: cost per task more than doubled versus Qwen3.7 Max even as token rates fell, and its GDPval-AA Elo lead over Kimi K3 came from taking about four times more turns per task. Watch: The same model regressed ten points on AA-Omniscience +14 to +4 with hallucination rate roughly doubling at flat accuracy — a model that stopped abstaining is materially riskier in unattended agent loops than its index score suggests. Method: Artificial Analysis patched the Intelligence Index to v4.1.1 with changed graders, so scores cited across index versions are no longer directly comparable — check the version before quoting a number in a design doc. Release: Cheaper defaults arrived from two directions: Muse Spark 1.2 puts Meta on the cost-per-task Pareto frontier, while open-weight Kimi K3 hit GA in Copilot and is callable through the LangSmith gateway — though GitHub has paused the K3 rollout while it mitigates an incident. Tooling: Agent Plugins landed as an open standard from OpenAI Devs and Cursor on the same day, letting a skill or MCP configuration authored once ship to Codex, Cursor, Copilot and VS Code without per-client repackaging. Watch: The agent-facing web stack advanced in parallel — a stateless next-generation MCP that runs as ordinary request handlers, WebMCP interfaces for existing sites, an agent-first browser in V8 isolates, managed retrieval via Cloudflare AI Search, and a protocol sketch for a readable, callable, payable agentic internet. Method: Two papers push back on summarization as a default context strategy: FinPerMA finds summarization drops the preference signal personalization depends on, with plain retrieval outperforming it, while SONAR argues code summaries for agent consumption should optimize correctness and abstraction level rather than polish. Release: OpenAI moved GPT-5.6 Sol behind all paid chats with a quantified factuality gain, expanded GPT-5.6 Luna to free users and removed free-tier chat limits, changing what you can assume about the model your users are actually on.