Good morning. Chip day, essentially. OpenAI unveiled benchmark numbers for its first custom silicon, Apple pushed to 2nm with the M6, and Qwen is about to drop a model sized precisely for the memory ceilings both machines advertise. If you squint, it’s the same story told three ways: inference is where the money and the engineering are going.
OpenAI’s Jalapeño chip posts strong first-generation numbers. At Hot Chips, OpenAI shared benchmarks for Jalapeño, its custom inference ASIC co-designed with Broadcom, claiming 1.5–1.9x better performance per watt and up to 3.6x lower latency versus Nvidia’s GB200/GB300 on several large models. The Verge and TechCrunch both note limited deployment by late 2026, scaling through 2027 — by which point Nvidia’s roadmap will have moved. SemiAnalysis points out the unusual part: first-gen custom silicon rarely beats incumbents, and the chip was designed in roughly 16 months. HN commenters compared the moment to the early 3dfx/PowerVR era and wondered aloud whether OpenAI could go further and bake specific model weights directly into silicon.
Apple’s M6 goes 2nm; M5 Ultra targets local inference. Apple announced the M6 and M5 Ultra, with the Ultra using a quad-die architecture, up to an 80-core GPU, and 1.2TB/s memory bandwidth. The new Mac Studio tops out at 512GB unified memory and can be clustered for distributed inference, with Apple explicitly pitching “local AI” above the fold. Pricing is the sore spot: a maxed 256GB Studio runs about $10K, and the 512GB config due in October is expected to clear $18K. Rumors also suggest Apple is skipping M6 Pro/Max/Ultra to concentrate on M7.
Qwen 3.8-Flash-Next arrives tomorrow, sized for that hardware. Alibaba’s Qwen team is releasing Qwen3.8-Flash-Next, a 125B-parameter MoE with 6B active parameters, framed as a preview of architectural work headed into the Qwen4 family. The 128GB memory footprint lines up almost too neatly with Strix Halo, GB10, and the mid-tier Mac Studio configs. One HN commenter estimated a 4-bit MLX quant with a 128K window running at 50–70 tokens/sec on a maxed-out MacBook Pro.
Multiverse Computing shows a 4-bit model beating its bf16 parent. In a Hugging Face writeup, Multiverse introduces Quantization-Aware Healing, applied to a GPT-OSS model compressed from 120B to 60B parameters and quantized to MXFP4. The recovered model beat the full-precision original on 7 of 9 benchmarks — smaller, cheaper, and slightly more accurate all at once. It’s another data point for the idea that inference costs have room to keep falling independent of new hardware.
Robotics and image-gen funding rounds land. Generalist, the robotics foundation-model startup staffed by DeepMind and Boston Dynamics alumni, hit a $3B valuation on a $200M Series B extension, with its Gen 1.5 model reportedly learning tasks from 3–12 second video demos. It’s still well behind Physical Intelligence ($11B) and Skild ($14B). Separately, Stability AI raised $76M from Universal, Sony, Warner, and EA — investors that also double as licensing partners, which is either clever alignment or a quiet admission about where the training data has to come from now.
That’s the morning. If Jalapeño’s numbers hold up under independent testing, the Nvidia margin conversation gets more interesting in a hurry.