The open weights frontier moves again, and every gain is quoted with its cost GLM-5.3 tied the leading open weights score with 60 on the Artificial Analysis Intelligence Index, matching Kimi K3, and posted a 246-point jump in agentic Elo from 1524 to 1770, second only to Opus 5, while burning about 20 percent more output tokens per task than GLM-5.2 at 68 cents per index task against 44. OpenAI halved GPT 5.6 Sol pricing on OpenRouter, Anthropic extended raised Claude Code limits through August 31, and IBM Research found that the useful amount of distilled guidance in agent memory is calibrated per model. Agents gained new capabilities: Claude can send Gmail messages, Cloudflare began per-request billing through x402, and Vercel posted $1 million against its sandbox to test isolation limits. GLM-5.3 tied the leading open weights score and posted a 246 point jump on agentic work, second only to Opus 5, while burning about 20 percent more output tokens per task than the model it replaces. The rest of the day read the same way: IBM measured the point past which more agent memory hurts, a retrieval study found recall-maximizing configurations resolving fewer issues under a fixed context budget, and OpenAI halved the price of Sol on one gateway while Anthropic held raised Claude Code limits open through the end of the month. Underneath that, the plumbing for agents that act rather than answer kept landing: send access in Gmail, per-request billing for agent traffic, and a million dollars posted against a sandbox to find out where it breaks. Read: GLM-5.3 scores 60 on the Artificial Analysis Intelligence Index, level with Kimi K3, and lifts its agentic Elo from 1524 to 1770, second only to Opus 5. The cost side moved with it: roughly 18,700 output tokens per task, about 20 percent more than GLM-5.2, at 68 cents per index task against 44. Weights are expected within the week. Read: Access terms shifted in four places. OpenAI cut GPT 5.6 Sol pricing by half on OpenRouter alone, Anthropic extended its raised weekly Claude Code limits through August 31 with a warning that capacity may be tight, Cowork opened on mobile and web for paid plans, and Glean made the case that routing, including answering some queries without a model at all, is now its own deployment layer. Read: Three separate results argue that more context is not free. IBM Research scaled its memory system across eight models and found the useful amount of distilled guidance is a dose calibrated per model rather than a switch to flip. A retrieval study found recall-maximizing configurations resolving fewer issues once the context budget is fixed. A controlled run on materials simulation measured what extra prompt detail actually buys when an agent writes domain-specific code. Read: Measurement moved into production. LangSmith shipped tuned evaluators that score live agent traces, claiming better accuracy than the frontier judges it tested against at 82 percent lower cost, and Artificial Analysis published a search index that puts answer quality and total task cost for agent search providers on one axis, with Parallel and Firecrawl on the frontier. Read: Agents got hands. Claude can now send Gmail messages and manage Drive files behind a user-set approval gate, Cloudflare began billing agents per request through x402 so MCP servers and API providers can charge for their own traffic, and Vercel put a million dollars against its sandbox to find where isolation for untrusted agent code gives way. Discuss: Both frontier labs published on pacing, from opposite ends. OpenAI paused reinforcement learning training on deployment-bound models for two weeks and is holding its largest planned run while it hardens and monitors its research environments, citing evidence of critical cyber capability. Anthropic reported Claude designing protein binders against 14 of 15 targets, with two outside labs building and testing them at binding rates above the published field baseline. Read: Recovery drew its own cluster of work: an open-weight repair agent reaching frontier-adjacent rates on locally hosted models, an evaluation of code repair once root causes span processes instead of one, an architecture for agent workflows that survive interruption and stay auditable, a reproduction showing how a retried tool call after a dropped connection can double-write, and Cursor on operating its Git storage as a database to keep repository reads reliable at scale.