# GPT-6.1 Sol: Pricing, Benchmarks and the Upgrade Case

> Source: <https://www.digitalapplied.com/blog/gpt-6-1-sol-pricing-benchmarks-upgrade-guide>
> Published: 2026-09-29 00:00:00+00:00

OpenAI’s GPT-6.1 Sol launch gives teams running AI agents a practical question: how much of their expensive model’s work can a cheaper model now finish well? The announcement positions Sol close to Astra on coding, computer use and professional work. That makes it a candidate for the everyday workload, with Astra retained where the extra capability earns its cost.

Our view: test Sol on work with a clear acceptance check first. A code change that passes review or a document answer supported by the right page is easier to evaluate than an open-ended strategy task. The business case depends on what survives that check.

1. 01Separate token rates from task costsA lower rate helps only if retries, longer answers and human review do not consume the saving.
2. 02Measure the cache you actually reuseAn agent repeatedly reads its instructions, tools and conversation. Distinguish cache hits, new writes and ordinary input in the bill.
3. 03Move a workload after it passesKeep the same tasks, permissions and acceptance checks during comparison. Change the model before changing the rest of the system.

## 01 — The releaseAvailable in Work, Codex and the API

According to OpenAI’s [September 29 announcement](https://openai.com/index/introducing-gpt-6-1-sol/), GPT-6.1 Sol is available to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, and through the API as `gpt-6.1-sol`. It is not yet available in Chat. Sol Ultrafast, promising up to eight times faster token generation in Codex, is announced for the coming days.

A planned speed option is a separate buying decision. Measure how long a complete job takes: tool execution, waiting for applications and review may occupy more time than generating the answer.

## 02 — The economicsSame input and output rates as Sol, cheaper cache reads

The current model pages for [GPT-6.1 Sol](https://developers.openai.com/api/docs/models/gpt-6.1-sol), [GPT-6 Sol](https://developers.openai.com/api/docs/models/gpt-6-sol) and [Astra](https://developers.openai.com/api/docs/models/gpt-6-astra) list the following standard API rates. A token is a unit of text the model processes; input includes what it reads, and output is what it generates.

| USD per million tokens, checked September 29, 2026. Standard processing at up to 272,000 input tokens per request. |  |  |  | 
|---|---|---|---|
| Billing category | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra | 
|---|---|---|---|
| Ordinary input | $2.00 | $2.00 | $10.00 | 
| Cached input | $0.10 | $0.20 | $1.00 | 
| Cache writes | $2.50 | $2.50 | $12.50 | 
| Output | $10.00 | $10.00 | $50.00 | 

Against Astra, Sol’s ordinary input and output rates are 80% lower; its cached-input rate is 90% lower. Against the original Sol, the price change is confined to cache reads. None of those percentages describes the saving on an entire agent run.

Above 272,000 input tokens, GPT-6.1 Sol’s model page specifies double input and cache rates and 1.5 times output rates for the full request. The large context window is therefore not a promise of the headline price at every prompt length.

At the standard rates, 100,000 cached tokens plus 10,000 ordinary input tokens cost $0.03 on GPT-6.1 Sol. Add 5,000 billable output tokens and the total is $0.08. The same mix costs $0.09 on GPT-6 Sol and $0.45 on Astra. This example assumes no cache writes and excludes tool fees, speed and regional premiums, retries and review.

OpenAI’s [prompt-caching guide](https://developers.openai.com/api/docs/guides/prompt-caching) distinguishes ordinary input, cached tokens and cache-write tokens. Reuse requires a matching prompt prefix and an eligible cache boundary. Shared text alone does not guarantee a hit.

Think of a support agent reading the same policy before each case. That stable material is a potential saving; each new case brings fresh input. If the application rewrites the earlier material, the saving can disappear. Compare the actual usage categories before and after switching, including the first write. Our [frontier-model API price index](https://www.digitalapplied.com/blog/frontier-model-api-price-index) provides the wider pricing context.

## 03 — The evidenceStronger results across several kinds of agent work

OpenAI reports these improvements over GPT-6 Sol in the launch evaluations:

- **DeepSWE v1.1, software engineering:** a 6.4-percentage-point gain over Sol’s best score, using lower effort and cost.
- **AutomationBench, business workflows:** a 4.8-point gain at medium effort.
- **OSWorld 2.0 offline set, computer use:** a seven-point gain at maximum effort for less than half the task cost; the metric is partial reward.
- **Terminal-Bench Science 0.1:** more than twice Sol’s score at maximum effort for less than half the task cost. Astra still leads the tested models.

These are OpenAI’s reported results, not Digital Applied tests. OpenAI notes that its research environment and API evaluations can differ from production ChatGPT. Different tools, prompts and reasoning settings affect the comparison.

A percentage-point improvement describes the gap between scores. Partial reward gives credit for progress within a task; it should not be read as the percentage of jobs completed perfectly. For a buyer, the useful signal is that Sol merits testing on more kinds of work. The numbers cannot tell you how often it will finish your particular process without intervention.

## 04 — The integrationCheck the request before swapping the model

The model reference and OpenAI’s [GPT-6 migration guide](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-6-astra) agree: GPT-6.1 Sol supports `low`, `medium`, `high`, `xhigh` and `max` effort, with medium the default. It does not support `none` or `minimal`. Tool calling requires the Responses API; Chat Completions supports requests without tools.

That matters for a team moving a non-reasoning Sol integration. A model-ID replacement alone may leave an unsupported effort or an unsuitable endpoint. Review the migration guidance, then test the requests your application sends, including failure handling.

For the comparison, hold the task set and tool access constant. Start at the effort level you intend to pay for. Record accepted results, total spend, elapsed time, retries and reviewer minutes. Keep difficult and ordinary tasks in separate groups: an average can hide a model that does well on routine cases and fails on the exceptions your team most needs it to handle.

## 05 — The operating limitsBetter alignment still needs enforceable permissions

OpenAI’s [system-card addendum](https://deploymentsafety.openai.com/gpt-6-1-sol) reports no attempts to bypass its automated reviewer in that evaluation. It also reports unwanted persistence in 23.5% of Sol rollouts in a warnings test, compared with Astra’s 17.4%. That test runs without the system controls designed to prevent circumvention; it is not a production failure rate.

The practical inference is to keep authority in the application. If a model can prepare a purchase but needs approval to submit it, enforce that distinction in the tool. A successful benchmark run should not expand what the agent can do with customer records, external messages or payments.

Include permission failures in the rollout test. A useful agent must explain a blocked action and complete the authorized work it can still do. That behaviour belongs in the acceptance criteria alongside correctness and cost.

## 06 — The buying decisionMove the tasks where the saving survives review

Our recommendation is a workload-level rollout. Choose a recurring task with a known finish line, compare accepted results and expand only where the evidence holds. The expensive model can remain available for exceptions without being the default for every step.

### Price an accepted result

GPT-6.1 Sol deserves a place in the evaluation queue. Its strongest business case will be a task your team already runs, finished to the same standard with a smaller total bill. Count the failed attempts and the reviewer’s time, then change the default where the saving remains.
