OpenAI adds GPT-6.1 Sol Ultrafast at six times standard API prices OpenAI added GPT-6.1 Sol to its Ultrafast service tier on October 8th, charging $12 per million input tokens and $60 per million output tokens in the API — six times the model's standard rates of $2 and $10. OpenAI's developer account said access is rolling out across the API, Codex and ChatGPT Work, with speeds up to 6x faster token generation in the API and up to 8x in Codex, while Codex and ChatGPT Work access is limited to Pro 500, eligible usage-based Enterprise and credit-based Edu plans. The tier completes a rollout OpenAI previewed on September 29th, and prompts exceeding 272,000 input tokens use different rates, with regional processing adding a 10% premium. OpenAI adds GPT-6.1 Sol Ultrafast at six times standard API prices The speed tier launched October 8th across the API, Codex and ChatGPT Work, with access limits for workplace plans and support for US and EU data residency. By Ryan Merket https://runtimewire.com/author/ryan-merket · Published · Updated Primary source: X https://x.com/OpenAIDevs/status/2108262812489531498 Why it matters OpenAI is making inference latency a premium line item: GPT-6.1 Sol Ultrafast costs six times standard API rates, pushing teams to decide which agent or interactive workloads merit faster output and which can wait. OpenAI added GPT-6.1 Sol https://runtimewire.com/models/openai/gpt-6.1-sol to its Ultrafast service tier on October 8th, charging $12 per million input tokens and $60 per million output tokens in the API, six times the model's standard rates. The company's developer account said in an October 8th thread https://x.com/OpenAIDevs/status/2108262812489531498 that access is rolling out across the API, Codex and ChatGPT Work. https://x.com/OpenAIDevs/status/2108262812489531498 https://x.com/OpenAIDevs/status/2108262812489531498 The release completes a rollout OpenAI previewed on September 29th, when it introduced GPT-6.1 Sol and said the faster mode would follow in the coming days. The launch announcement https://openai.com/index/introducing-gpt-6-1-sol/ positioned Sol as a lower-cost model for complex coding and professional work, with standard API rates of $2 per million input tokens and $10 per million output tokens. Ultrafast keeps the same model and sells a shorter wait for its output. OpenAI's stated speed ceilings | Product | Maximum speed claim | |---|---| | API | Up to 6x faster token generation than Sol Standard | | Codex | Up to 8x faster token generation than Sol Standard | The October 8th thread describes speeds up to eight times faster than Sol Standard. The DevDay recap https://openai.com/index/devday-2026-recap/ specifies up to eight times faster token generation in Codex and up to six times in the API. The company's Ultrafast documentation https://developers.openai.com/api/docs/guides/ultrafast-mode calls it the API's fastest service tier and recommends it when speed justifies the higher cost. These are OpenAI's stated maximums, not guaranteed response times for every request. For developers, the price gap is large enough to affect which calls belong on the faster tier. The short-context API rates per million tokens are: | Token type | Standard | Ultrafast | |---|---|---| | Input | $2 | $12 | | Cached input | - | $0.60 | | Cache writes | - | $15 | | Output | $10 | $60 | At those standard input and output rates, one million input tokens plus one million output tokens costs $12. The same volume costs $72 in Ultrafast, before any other charges. OpenAI's pricing page https://developers.openai.com/api/docs/pricing lists the corresponding short-context Ultrafast prices for GPT-6.1 Sol. The model page https://developers.openai.com/api/docs/models/gpt-6.1-sol says prompts exceeding 272,000 input tokens use different rates, and regional processing adds a 10% premium where available. OpenAI is selling speed as a selectable operating cost for work where developers say latency has consequences: the thread names outage debugging, agents navigating applications and live experiences. A team can reserve the six-times-priced tier for latency-sensitive requests and leave other traffic on standard service; whether that improves a particular workflow enough to justify the bill depends on its own tests. The thread's examples describe intended use cases, not independently measured customer results. The tier's interface is also unevenly gated. OpenAI says API access is available to all users, while access in Codex and ChatGPT Work is limited to Pro 500, eligible usage-based Enterprise and credit-based Edu plans; Enterprise administrators must enable it. The API guide now lists GPT-6.1 Sol Ultrafast rate limits of 1 million tokens per minute for Build, 4 million for Launch and 40 million for Grow. These are separate from Standard and Fast limits, so a faster mode does not remove throughput constraints. OpenAI recommends persistent WebSocket connections for agentic applications that make successive tool calls, warning that network overhead can eat into the latency benefit. Its documentation confirms that GPT-6.1 Sol Ultrafast supports US and EU data residency, alongside global processing. That broadens the service's fit for some organizations with residency requirements, while the premium remains attached to token usage. Sam Altman, OpenAI's co-founder and CEO, helped frame the company at its 2015 launch around making AI broadly distributed. The current product strategy puts that idea into a tiered commercial form: OpenAI offers a lower-priced model for complex work, then charges a substantial premium when customers want its output sooner. For developers building interactive agents, the sales pitch is less about a new model capability than buying back time on the calls where users notice the wait.