Navigating the solution space with agents
Coding agents dramatically lower the cost of exploring solution spaces but remain path-dependent and poor at deciding optimal endpoints, argues Mitchell Hashimoto and others in a debate about reading …
Coding agents dramatically lower the cost of exploring solution spaces but remain path-dependent and poor at deciding optimal endpoints, argues Mitchell Hashimoto and others in a debate about reading …
OpenAI's GPT-5.6 model family, previewed in late June to government-vetted partners and fully released on July 9, includes three variants—Sol, Terra, and a third—with a 1.05-million-token context wind…
Nearly four years after ChatGPT's release, large language models still suffer from the same core flaws — poor math, outdated knowledge, short memory, and toxicity — that were present in GPT-3.5, accor…
OpenAI has merged its ChatGPT, Work, and Codex applications into a single unified interface, with the latest Mac app update consolidating the previously separate tools under the ChatGPT brand. The red…
Ben Thompson of Stratechery argues that China's AI strategy is to commoditize complements, leveraging its lead in robotics and physical-world dominance through widely available AI models, while weaken…
OpenAI's Sol model autoformalized a ChatGPT-disproved Erdős Unit Distance conjecture in Lean, generating 1.2 million lines of code in three weeks. Boris Alexeev of OpenAI steered the model to a comple…
Kimi K3, developed by Moonshot, achieved top-tier performance on a private cybersecurity benchmark, offering strong recall, precision, and cost efficiency, while GPT 5.6 leads in recall and precision …
Claude users report that the Fable model has been removed from their accounts without warning, requiring credits to continue usage, disrupting ongoing development work. Multiple users on Hacker News s…
OpenAI proposed a new AI scorecard called 'Useful Intelligence per Dollar' on July 17, 2026, designed to help enterprises measure AI value through four lenses: useful work produced, cost per successfu…
OpenAI released GPT-5.6 to the public on July 9 as three tiers — Sol (flagship), Terra (balanced), and Luna (budget) — all running on the same 4 trillion parameter Spud pretrain base, with pricing ran…
Chinese President Xi Jinping, in a speech marking 70 years since the Dartmouth workshop, called for global cooperation to ensure AI is developed for positive and good purposes, emphasizing a people-ce…
Kimi K3 ranks third on Artificial Analysis's intelligence index, trailing the leader by only two points. The independent evaluation platform assesses leading AI models across intelligence, cost, speed…
The U.S. government's process for evaluating advanced AI models like OpenAI's Sol and Anthropic's Fable lacks transparency, according to Senior Research Analyst Mina Narayanan, who said she has no vis…
OpenAI's GPT-5.6 Sol is the best vision model the company has released, scoring 46.2 mAP@50 in object detection compared to GPT-5.5's 13.8, according to Roboflow's upcoming VLM benchmark. Sol also ach…
OpenAI shipped GPT-5.6 on July 9 as three distinct models — Sol, Terra, and Luna — not a single upgrade, with the gpt-5.6 alias routing to Sol at $5 per million input tokens. The release introduces Pr…
OpenAI introduced GPT-Red on July 15th, an internal automated red-teaming model designed to find prompt injection vulnerabilities at scale and feed those attacks back into the training of production m…
Since early June, OpenAI's coding tool Codex encrypts the instructions a main agent passes to its subagents, leaving developers unable to track internal task delegation. For the larger GPT-5.6 variant…
OpenAI released three models named Luna, Terra, and Sol, which correspond to small, medium, and large variants of GPT-5.6. The naming is not branding but an inference cost architecture announcement th…
OpenAI has removed the five-hour Codex cap for some accounts, replacing it with a weekly limit that better suits long coding sessions, according to a user report. The same source also highlights the r…
OpenAI released GPT-5.6 on July 9, 2026, as a three-tier model family—Sol, Terra, and Luna—all optimized for agentic tool calling. Each tier shares a 1M-token context window, 128K max output, and nati…