The Roundup No. 1 OpenAI shipped the GPT-5.6 family — Sol, Terra, and Luna — with pricing from $1 to $5 per million input tokens and 1M context, partly served on Cerebras hardware, making frontier-quality tokens cheap enough to waste. Google reportedly delayed Gemini 3.5 Pro by months due to shortfalls in coding and reasoning, while Moonshot's Kimi K3 went open-weight and took #1 on Frontend Code Arena with a 76% win rate. xAI released Grok 4.5, a 1.5T-parameter coding-focused mixture-of-experts model at $2/$6 per million tokens, and other releases included Gemini 3.6 Flash, Poolside Laguna S 2.1, GPT-Live, Cognition SWE-1.7, and Meta's Muse Image, Muse Video, and Muse Spark 1.1. What actually mattered: Jul 7–21 Two weeks, five stories that matter, a handful that don’t need more than a sentence. Four minutes, in the order it matters. OpenAI ships the GPT-5.6 family — and the pricing is the story Sol is the flagship with a new Ultra subagent mode, Terra delivers GPT-5.5-level quality at half the cost, and Luna is the fast tier. $1–$5 per million input tokens, 1M context, partly served on Cerebras hardware. Why it matters Frontier-quality tokens just got cheap enough to waste. When capability stops being the constraint, workflow design becomes the whole game — which is exactly where most teams are furthest behind. Read my full day-one review → /gpt-5-6-review Google reportedly delays Gemini 3.5 Pro by months Internal testing showed shortfalls in coding and long-horizon reasoning; it stays in limited enterprise preview. The two-horse frontier race just got lonelier at the front. Kimi K3 goes open-weight and takes 1 on Frontend Code Arena Moonshot’s model wins 76% of matchups. The open-weight frontier keeps compressing the gap on exactly the tasks that used to justify closed-model pricing. xAI releases Grok 4.5 A 1.5T-parameter coding-focused mixture-of-experts at $2/$6 per million tokens — aggressive pricing aimed squarely at the same developers OpenAI courted two days later. Gemini 3.6 Flash lands — cheaper Flash tier; ~17% fewer output tokens on agentic workloads. Poolside releases Laguna S 2.1 — a 118B open-weight model for agentic coding. GPT-Live launches — full-duplex voice for ChatGPT; it listens and speaks simultaneously. Cognition ships SWE-1.7 — an RL-tuned coding model running at ~1,000 tokens/sec. Meta ships Muse Image and Muse Video , plus Muse Spark 1.1 with 1M context two days later. Four minutes, a few times a week. Never a wasted send. If a week is boring, you don’t hear from me.