{"slug": "what-216-coding-attempts-revealed-about-sol-6-1-s-value", "title": "What 216 Coding Attempts Revealed About Sol 6.1's Value", "summary": "A developer's Tor Production benchmark of 216 coding attempts across Sol 5.6, Sol 6, and Sol 6.1 found that Sol 6.1 at Low effort averaged $0.101 per attempt with a perfect 100/100 main score and 12/12 accepted deliveries, the cheapest fully observed setting to achieve that. On a transaction task, Sol 6.1's added tests caught 4/6 and 5/6 seeded defects at Low effort, rising to 5/6 and 6/6 at Medium for 9.9% more estimated cost, while its API-equivalent total ran 28.9% below Sol 6 across 70 matched observations per model.", "body_md": "**TL;DR**\n\n- Sol 6.1 Low averaged\n**$0.101 per attempt**, scored **100/100** on the main checks, and delivered accepted code in **12/12 attempts**.- On the transaction task, its added tests caught\n**4/6 and 5/6 seeded defects**. Medium improved that to **5/6 and 6/6** at 9.9% more estimated cost.- These are results from six specified tasks. Dollars are\n**frozen API-equivalent estimates**.\n\nI compared Sol 5.6, Sol 6, and Sol 6.1 in a benchmark published through my Tor Production project. The most useful differences appeared when I looked beyond passing scores: **what did the generated tests catch, and what did that capability cost?**\n\nThe [public repository](https://github.com/Tor-Production/sol-coding-benchmark) contains the stand, task contracts, evaluators, submitted code, results, and English report.\n\nThe setup used **three models × six reasoning efforts × six Python tasks × two attempts = 216 attempts**. The efforts were Low, Medium, High, Xhigh, Max, and Ultra.\n\nThe tasks covered:\n\nEach attempt started with a fresh scaffold and session. Agents could run supplied tests and add tests, but received no hidden-test feedback or grading retry. The archive retains failed outcomes: **214 attempts completed; 213 were accepted deliveries**.\n\nSixteen of the 18 model-and-effort settings reached **100/100** on the main functional score. That ceiling leaves little room for a useful quality ranking from the main score alone.\n\nAt Low effort, the comparison looked like this:\n\n| Model | Mean estimate / attempt | Main score | Accepted deliveries | \n|---|---|---|---|\n| Sol 5.6 | $0.318 | 100/100 | 12/12 | \n| Sol 6 | $0.125 | 99.44/100 | 11/12 | \n| Sol 6.1 | $0.101 | 100/100 | 12/12 | \n\n**Sol 6.1 Low was the cheapest fully observed setting with a perfect main score and all 12 deliveries accepted.**\n\n*The chart covers all efforts; the table isolates Low.*\n\nA saved implementation's grade and a completed delivery are separate outcomes. For example, a Sol 6 Max reservation attempt timed out although its saved code passed the main checks. It reduced the accepted-delivery rate, while retaining its passing functional score.\n\nAll **36 transaction implementations** passed the main evaluator. Their added tests were much less alike.\n\nI evaluated those tests against six fixed defects injected into a correct reference implementation: snapshot-read errors, write skew, missed phantoms, forgotten reads after rollback, recovery-version gaps, and checkpoint aliasing.\n\n| Model, Low effort | Mean estimate / MVCC attempt | Defects caught: run 1 | Run 2 | \n|---|---|---|---|\n| Sol 5.6 | $0.305 | 0/6 | 0/6 | \n| Sol 6 | $0.137 | 3/6 | 0/6 | \n| Sol 6.1 | $0.129 | 4/6 | 5/6 | \n\nThe **three zero results** mean those attempts had no recognized candidate-added tests. The implementations still passed their main checks; the zeros apply to this test-strength criterion.\n\nA usable mutation result requires the added tests to pass on both the candidate and the positive reference. Only executed assertion failures count as detected defects. Incompatible or incomplete evidence is reported as **N/A**, rather than averaged into a setting's score.\n\nFor Sol 6.1:\n\nThis makes the effort choice concrete: how much additional test sensitivity is useful for your task? The result applies to this declared defect set; it does not establish that higher effort always catches more real-world bugs.\n\nAcross **70 matched task/effort/repetition observations per model**, Sol 6.1's API-equivalent total was **28.9% lower than Sol 6's**. The same two partial-cost keys were excluded from both models.\n\nThe pricing bridge shows three steps:\n\nThe middle step changes the cached-input tariff while holding usage constant. The final step reflects the observed usage difference under a common tariff. This order-dependent comparison does not provide causal proof of intrinsic model efficiency.\n\n**Sol 6.1 showed strong value in this suite, especially at Low.** For the specified transaction task, Medium bought a measurable improvement in generated-test sensitivity at a modest estimated cost increase.\n\nWhen passing scores reach a ceiling, inspect another outcome that matters to your work. Here, tests that detect declared defects offered more useful separation than the aggregate main score.\n\nKeep the scope in view:\n\nThe [findings and derivations](https://github.com/Tor-Production/sol-coding-benchmark/blob/main/docs/findings.md) make every comparison inspectable.\n\n**What task or defect would you add to make the next comparison more useful?**", "url": "https://wpnews.pro/news/what-216-coding-attempts-revealed-about-sol-6-1-s-value", "canonical_source": "https://dev.to/yuriitor/what-216-coding-attempts-revealed-about-sol-61s-value-2noo", "published_at": "2026-10-06 18:38:30+00:00", "updated_at": "2026-10-06 18:48:34.749854+00:00", "lang": "en", "topics": ["ai-research", "large-language-models", "ai-tools", "developer-tools"], "entities": ["Sol 6.1", "Sol 6", "Sol 5.6", "Tor Production"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-216-coding-attempts-revealed-about-sol-6-1-s-value", "markdown": "https://wpnews.pro/news/what-216-coding-attempts-revealed-about-sol-6-1-s-value.md", "text": "https://wpnews.pro/news/what-216-coding-attempts-revealed-about-sol-6-1-s-value.txt", "jsonld": "https://wpnews.pro/news/what-216-coding-attempts-revealed-about-sol-6-1-s-value.jsonld"}}