# GPT-5.6 Sol clarification vs guess benchmark

> Source: <https://gist.github.com/ray2201/79db4a7b6b4e135513b3dbc1766131e4>
> Published: 2026-10-07 13:40:45+00:00

Question:

When is it cheaper to ask one clarifying question than let the model guess?

Model: GPT-5.6 Sol

Reasoning effort: Low

Tasks: 40

Ambiguity levels: 1 to 4 plausible choices

Two strategies were compared:

The model completes the task without the missing user preference.

If the first answer is wrong, a simulated user correction is provided and the model gets a second API call.

The system asks one deterministic clarification question first.

The missing preference is then supplied before the model completes the task.

Guess first:

- First-pass success: 75%
- Final success after correction: 100%
- Rework rate: 25%
- Average API calls/task: 1.25
- Average cost/task: $0.000771
- Median model latency: 1660 ms

Clarify first:

- Success: 100%
- Average API calls/task: 1.00
- Average cost/task: $0.000543
- Median model latency: 1458 ms

Clarification reduced API cost to successful completion by 29.5%.

1 choice:

- Guess success: 100%
- Clarification saving: 6.5%

2 choices:

- Guess success: 80%
- Clarification saving: 20.4%

3 choices:

- Guess success: 70%
- Clarification saving: 33.5%

4 choices:

- Guess success: 50%
- Clarification saving: 45.2%

Clarification adds one user interaction.

Human response time is not included in the latency numbers. latency figures are model/API latency only.

This is a small synthetic benchmark.
