Together AI's Zain Hasan tells LDS that its 83% coding-task success result was reconstructed from benchmark trials using hidden tests, rather than measured in a live router. The $3.35 figure covers inference per attempted task, or about $4.04 per solved task, before verification, sandbox time and human review. His answers explain what teams need to test before assuming that a cheaper first model will produce cheaper accepted code.
An AI coding workflow can look inexpensive until someone has to decide whether its patch is actually good enough to use. That decision is central to the economics of sending a task to a cheaper model first and escalating failures to a more expensive one.
Together AI's DeepSWE analysis reports an 83% success rate at $3.35 per task for a DeepSeek-first, GPT-second cascade. In an original written interview with Let's Data Science, Zain Hasan, Staff AI/ML Engineer at Together AI, clarifies what that result measures, which costs it leaves out and why a real repository cannot simply copy the benchmark's decision rule.
The important engineering problem is not just generating another patch. It is recognizing when the first patch should be accepted, retried or escalated.
The cascade was reconstructed, not run as a live router
Hasan's first clarification changes how the headline result should be read:
"We reconstructed the cascade from the official pass@4 trials rather than running it live."
In the reconstruction, DeepSeek's attempt was checked against held-out tests. If it failed even one, GPT received the same problem afresh. It did not receive DeepSeek's patch, test output or feedback.
That is a second independent attempt, not a measured workflow in which one model diagnoses and repairs another model's work. The models were deepseek-v4-pro-0813 and gpt-5.6-sol, both at the reported max setting. The underlying comparison used 113 tasks and four trials per model, totaling 904 rollouts.
The held-out tests supplied an exact acceptance signal for the benchmark. A real team's tests may be incomplete, and passing them does not establish that a patch satisfies requirements they do not cover. The 83% result therefore describes success under the benchmark's scoring conditions, not an observed production acceptance rate.
A passing patch can still make the application worse
Hasan illustrates the problem with a caching example. Users do not immediately see changes to their profiles. Instead of fixing cache invalidation, the agent removes the cache altogether. The visible bug disappears and the functional tests pass, but every page load now queries the database.
This is a teaching example supplied in his answer, not a failure observed in the benchmark or a named customer incident. It shows why checking the reported symptom may miss a performance regression.
A performance check or a reviewer inspecting the change could identify the problem. Hasan recommends adding newly discovered failure cases to the team's evaluation over time, so the checks better reflect the work the application actually does.
For production code, he proposes tests written before the coding attempt, supplemented by regression tests. Teams also need checks relevant to the change, including security and performance where applicable. Those requirements apply to outputs from both models. Escalating to a more capable model does not remove the need to review its work. Model-based judges and scoring systems can help select candidates, Hasan says, but their accuracy and speed need evaluation too. A verifier that approves the wrong patch undermines the routing decision; a slow one can become the bottleneck.
Count more than the model bill
Hasan confirms that $3.35 is average inference cost per attempted task across the reconstructed cascade. With 83% of tasks solved under its scoring conditions, that becomes approximately $4.04 per solved task: $3.35 divided by 0.83. He confirmed that calculation in a follow-up to LDS.
He also states the boundary of the cost estimate:
"These figures cover only inference, meaning the token costs of running the models."
Sandbox time, verification, human review and the cost of waiting for a fix are excluded. A benchmark-solved task is also not necessarily a patch accepted into a real codebase.
His illustrative cost breakdown shows why this matters. Reviewing 83 patches for 15 minutes each at an assumed $60 per hour costs $1,245. The review duration and hourly rate are assumptions, not measured staffing costs; Hasan confirmed the arithmetic after LDS queried the example.
That does not establish that every team will spend more on review than inference. It demonstrates how labor can outweigh a low token bill under explicit assumptions. Teams need to measure their own review burden, including time spent rejecting or repairing incorrect patches.
Urgency adds another cost. A cheaper first attempt can delay an expensive second attempt while a person or service waits. Hasan suggests going directly to the more capable model for urgent work where that delay matters. Whether it helps on a particular workload remains a question for the team's own evaluation.
Compare routing with retries, not just one expensive attempt
The comparison changes when a single model is allowed another try. In his response, Hasan gives the following success rates under the held-out-test scoring:
- •GPT-5.6 Sol with up to two attempts: 81.0%.
- •DeepSeek V4 Pro 0813 with up to two attempts: 78.5%.
- •DeepSeek followed by GPT on failure: 83.0%.
- •DeepSeek with up to four attempts: 88.5%.
These figures describe finding at least one answer that passes the benchmark tests. They do not show that a production system can reliably choose that answer from several candidates.
Four attempts can improve coverage while adding waiting time or concurrent work. Retrying and switching models also consume different budgets. Hasan explicitly separates the reported analysis from comparisons it did not perform:
"We did not measure matched-budget comparisons or selection with realistic checks."
The task sample also limits interpretation. Hasan notes that DeepSeek's four-attempt coverage corresponds to 100 tasks, versus 97 for Sol. A difference of three tasks on this sample is not evidence that the same ranking will hold on another company's repositories.
The useful comparison for an engineering team is consequently broader: a cascade, retries of each model alone, and the actual checks available to choose or reject their outputs. Measure them under comparable spending and time constraints before deciding which workflow earns its place.
A small repository evaluation that tests the acceptance decision
Asked for a runnable example, Hasan says the team has not built one. He points to the published trial records and offers a proposed evaluation approach. The following is a plan derived from his answer, not a tutorial LDS has executed or a reproduction of the reported score.
Start with roughly 20 to 30 closed issues that a person fixed and that have associated tests. Restore each repository to its state before the fix, give the model the issue description and reserve the human-written test for final scoring. Hasan cautions that public repository history may already have appeared in training data.
Keep two judgments separate. First, use the checks available to the workflow to decide whether a candidate can go forward for review. Second, use the withheld scoring test to assess whether that decision was right. Feeding the withheld result back into the router would restore the benchmark advantage the exercise is meant to examine.
His proposed checks include whether the patch applies, the existing and prewritten tests, linting and a security scan. These are components of a review process, not a guarantee of correctness or security. The caching example explains why task-specific checks may still be necessary.
Record the evidence needed to make a decision:
- •How often the workflow accepts a patch that later fails the withheld check.
- •The total cost per correct accepted patch, including unsuccessful attempts and review.
- •Elapsed time until an acceptable result, not just model response time.
- •How often escalation produces a correct answer after the cheaper model fails.
- •Whether retries of one model achieve comparable results under the same constraints.
Security findings, changes to protected evaluation tests and exhausted time or cost limits should trigger review rather than automatic acceptance. Appropriate thresholds depend on the repository and the consequences of a mistake.
For teams building coding agents, this is the strongest practical lesson in Hasan's answers: the inexpensive model and the acceptance checks must be evaluated together. A low generation price only becomes useful savings when the workflow can distinguish a suitable patch from one that merely looks successful.
Reporting note
This LDS Exclusive is based on five original written answers and follow-up cost and attribution confirmations supplied on September 29, 2026. Responses are attributed to Zain Hasan, Staff AI/ML Engineer at Together AI. The company's DeepSWE analysis and linked trial records provide background. LDS has not rerun the benchmark or tested a production router. The caching scenario, cost assumptions and repository exercise are illustrative; the reported cascade uses held-out-test selection.
Key Points #
- 1The 83% cascade result was reconstructed from separate trials using held-out tests. It was not measured in a live router with production acceptance checks.
- 2The $3.35 figure is inference cost per attempted task, about $4.04 per solved task. Verification, sandbox time, human review and waiting are excluded.
- 3Compare routing against single-model retries under comparable cost and time limits. Finding one correct candidate is different from reliably identifying and accepting it.
Scoring Rationale #
Original written interview clarifies benchmark reconstruction, inference denominators, excluded costs and realistic verification. Useful guidance is separated from measured evidence.
Sources #
Original reporting, with the public references used alongside it.
LDS Exclusive
Reporting based on written answers given directly to Let's Data Science by Zain Hasan, Staff AI/ML Engineer, Together AI.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.