TL;DR
- Sol 6.1 Low averaged $0.101 per attempt, scored 100/100 on the main checks, and delivered accepted code in 12/12 attempts.- On the transaction task, its added tests caught 4/6 and 5/6 seeded defects. Medium improved that to 5/6 and 6/6 at 9.9% more estimated cost.- These are results from six specified tasks. Dollars are frozen API-equivalent estimates.
I compared Sol 5.6, Sol 6, and Sol 6.1 in a benchmark published through my Tor Production project. The most useful differences appeared when I looked beyond passing scores: what did the generated tests catch, and what did that capability cost?
The public repository contains the stand, task contracts, evaluators, submitted code, results, and English report.
The setup used three models × six reasoning efforts × six Python tasks × two attempts = 216 attempts. The efforts were Low, Medium, High, Xhigh, Max, and Ultra.
The tasks covered:
Each attempt started with a fresh scaffold and session. Agents could run supplied tests and add tests, but received no hidden-test feedback or grading retry. The archive retains failed outcomes: 214 attempts completed; 213 were accepted deliveries.
Sixteen of the 18 model-and-effort settings reached 100/100 on the main functional score. That ceiling leaves little room for a useful quality ranking from the main score alone.
At Low effort, the comparison looked like this:
| Model | Mean estimate / attempt | Main score | Accepted deliveries |
|---|---|---|---|
| Sol 5.6 | $0.318 | 100/100 | 12/12 |
| Sol 6 | $0.125 | 99.44/100 | 11/12 |
| Sol 6.1 | $0.101 | 100/100 | 12/12 |
Sol 6.1 Low was the cheapest fully observed setting with a perfect main score and all 12 deliveries accepted.
The chart covers all efforts; the table isolates Low.
A saved implementation's grade and a completed delivery are separate outcomes. For example, a Sol 6 Max reservation attempt timed out although its saved code passed the main checks. It reduced the accepted-delivery rate, while retaining its passing functional score.
All 36 transaction implementations passed the main evaluator. Their added tests were much less alike.
I evaluated those tests against six fixed defects injected into a correct reference implementation: snapshot-read errors, write skew, missed phantoms, forgotten reads after rollback, recovery-version gaps, and checkpoint aliasing.
| Model, Low effort | Mean estimate / MVCC attempt | Defects caught: run 1 | Run 2 |
|---|---|---|---|
| Sol 5.6 | $0.305 | 0/6 | 0/6 |
| Sol 6 | $0.137 | 3/6 | 0/6 |
| Sol 6.1 | $0.129 | 4/6 | 5/6 |
The three zero results mean those attempts had no recognized candidate-added tests. The implementations still passed their main checks; the zeros apply to this test-strength criterion.
A usable mutation result requires the added tests to pass on both the candidate and the positive reference. Only executed assertion failures count as detected defects. Incompatible or incomplete evidence is reported as N/A, rather than averaged into a setting's score.
For Sol 6.1: This makes the effort choice concrete: how much additional test sensitivity is useful for your task? The result applies to this declared defect set; it does not establish that higher effort always catches more real-world bugs.
Across 70 matched task/effort/repetition observations per model, Sol 6.1's API-equivalent total was 28.9% lower than Sol 6's. The same two partial-cost keys were excluded from both models.
The pricing bridge shows three steps:
The middle step changes the cached-input tariff while holding usage constant. The final step reflects the observed usage difference under a common tariff. This order-dependent comparison does not provide causal proof of intrinsic model efficiency.
Sol 6.1 showed strong value in this suite, especially at Low. For the specified transaction task, Medium bought a measurable improvement in generated-test sensitivity at a modest estimated cost increase.
When passing scores reach a ceiling, inspect another outcome that matters to your work. Here, tests that detect declared defects offered more useful separation than the aggregate main score.
Keep the scope in view:
The findings and derivations make every comparison inspectable. What task or defect would you add to make the next comparison more useful?