OpenAI improves prompt caching in GPT-6 Sol and Luna for better efficiency OpenAI launched GPT-6 Sol and Luna on September 22, 2026, with prompt caching upgrades that cut cached input token costs by up to 90% and reduce latency for developers building agentic workflows. The new system raises default cache hit rates, adds diagnostic tools and explicit cache breakpoints, and lets developers toggle reasoning effort levels and swap tools without invalidating cached context, replacing the automatic prefix matching for prompts over 1,024 tokens introduced with GPT-4o. OpenAI priced GPT-6 API access at roughly 50% below GPT-5 promotional rates, and the caching improvements extend across the GPT-6 API, ChatGPT Work, and Codex, while flagship variant GPT-6 Astra entered limited preview on September 3, 2026, with general availability the next day. FoxTPNL / Wikimedia Commons CC BY 4.0 OpenAI improves prompt caching in GPT-6 Sol and Luna for better efficiency New caching upgrades in GPT-6 Sol and Luna cut cached input token costs by up to 90% while reducing latency for developers building agentic workflows. OpenAI launched GPT-6 Sol and Luna on September 22, 2026, with a suite of prompt caching upgrades that slash costs on cached input tokens by up to 90% and reduce latency for developers working with long conversational contexts. What actually changed under the hood The GPT-6 caching overhaul introduces several concrete upgrades. Default cache hit rates are higher out of the box, meaning developers get the cost savings without needing to manually optimize their prompt structures. New diagnostic tools let builders monitor exactly how well their caching is performing. Explicit breakpoints give developers fine-grained control over which portions of a prompt get cached and where those cached segments begin and end. Previously, OpenAI’s caching relied on automatic prefix matching for prompts exceeding 1,024 tokens, a system introduced with GPT-4o. The new system also lets developers toggle reasoning effort levels and swap available tools without invalidating the cached context. The pricing math and competitive implications OpenAI is pricing GPT-6 API access at roughly 50% below GPT-5 promotional rates. Combined with the 90% discount on cached input tokens, the effective cost of running sustained AI conversations or multi-step agent pipelines drops dramatically. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. GitHub Copilot previously achieved a reduction of over 50% in prompt tokens requiring fresh processing across billions of requests using earlier versions of OpenAI’s caching system. GPT-6 Astra and the broader rollout GPT-6 Astra, the flagship variant of the series, entered limited preview on September 3, 2026, with general availability following the next day. The caching improvements extend across multiple OpenAI products, including the GPT-6 API, ChatGPT Work, and Codex. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .