The harness is becoming the product OpenAI's GPT-5.6 Sol achieved a 38.3% ARC-AGI-3 score with a harness that retained reasoning and used compaction, up from 13.3%, while cutting output tokens sixfold. The same model helped OpenAI reduce serving costs by 20% and improve token-generation efficiency by over 15%. Cursor launched on iPad, and Similarweb introduced a deep-research evaluation stack, highlighting that system components like memory, orchestration, and verification increasingly determine AI capability. The strongest signal recovered from X today is not another model launch. OpenAI showed that GPT-5.6 Sol’s ARC-AGI-3 score moved from 13.3% to 38.3% when the harness retained reasoning and replaced rolling truncation with compaction—while output tokens fell sixfold. In a separate engineering account, the same model helped analyze production traffic, rewrite kernels and tune its own speculator, contributing to a 20% reduction in serving costs and more than 15% higher token-generation efficiency. Same model, radically different system. That systems shift is reaching the product edge. Cursor’s move onto iPad turns coding into a monitor, steer and review loop that can follow developers away from their desks. Similarweb’s deep-research evaluation stack—deterministic tool checks, rubric judges, faithfulness tests and saved baselines wired to traces—shows the operating discipline required once those agents are everywhere. Today’s news is less that models became smarter than that memory, orchestration, interfaces and verification increasingly decide what their intelligence can accomplish. Harness: Retained reasoning and compaction tripled GPT-5.6 Sol’s ARC-AGI-3 score and cut output tokens sixfold, demonstrating that an eval measures the surrounding system as much as the model. Infrastructure: GPT-5.6 Sol helped optimize OpenAI’s own routing, kernels and speculative decoding; the reported gains include 20% lower serving cost and more than 15% better token-generation efficiency. Interface: Cursor on iPad makes agentic coding look more like an asynchronous inbox and review surface than a desktop-bound editor. Evaluation: Similarweb’s deep-research recipe combines deterministic behavior checks, LLM quality rubrics, faithfulness and trace-based baselines instead of pretending one score can cover an open-ended task.