cd /news/artificial-intelligence/sudo-l7-a-benchmark-that-measures-th… · home › topics › artificial-intelligence › article
[ARTICLE · art-147816] src=surgehq.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Sudo L7 – A benchmark that measures the judgment behind good engineering

A new benchmark called sudo L7, built by an unnamed team, measures staff-level software engineering judgment across 60 tasks authored by professional software engineers, most drawn from private production repositories at real companies. The strongest agents tested, Claude Opus 5.5 and Sonnet 5.5, each completed roughly 45% of the tasks, which span backend, frontend, security, reliability, and architecture work in Python, TypeScript, JavaScript, and Ruby. The benchmark grades agents on eight dimensions of engineering behavior using expert-authored rubrics, including architectural judgment, verification, and communication, rather than only whether code passes tests.

read11 min views1 publishedOct 8, 2026
Sudo L7 – A benchmark that measures the judgment behind good engineering
Image: source

Appendix

The code works. Would you ship it? #

The strange thing about today’s coding agents is that they can be technically right and professionally wrong.

  • A password manager can implement every requested feature, while displaying saved passwords in plaintext.
  • A production database migration can correctly change a schema, while deleting records without confirmation and no way to roll back if something goes wrong.
  • A stock-alert dashboard can look finished, without notifying the user it’s simulating every price.

Models know how to write code. What they often miss is everything around it: what could go wrong, what the user forgot, what needs to be tested, and what another engineer would want to know before saying LGTM.

A junior engineer is often handed a well-defined problem, and asked for a patch. They follow tickets, write code, and check whether tests pass.

A staff engineer is handed something fuzzier: make this faster; migrate this safely; figure out what we should do. They notice what tickets forget to mention, anticipate consequences, and know when passing tests isn’t the end of the story.

That distinction is what we built sudo L7 to measure.

Introducing sudo L7 #

sudo L7 is our benchmark for staff-level software engineering: writing code, but also the professional judgment and behavior around it.

It contains 60 tasks authored by professional software engineers. Tasks are primarily grounded in private production repositories from real companies, alongside open-source projects. One repository, for example, comes from a money-management app and contains roughly 250,000 lines of code.

The requests look like things engineers actually get:

  • “Make this faster.”
  • “Migrate these API tokens to secure hashes.”
  • “Can we safely cache this?”
  • “We have a security audit coming up. Fix our password protection.”

Agents have to explore the system, decide what matters, choose an approach, and verify the dangerous cases.

Most coding benchmarks measure the inner loop of software engineering: given a well-specified problem, can the agent write code that passes? sudo L7 measures the outer loop as well: can the agent notice what to build, what not to build, where the risk is, and when “the tests pass” isn’t the end of the story?

Sometimes the output is code. Sometimes it’s architectural analysis, a rollout plan, or an explanation of why the obvious implementation creates a larger problem.

All that is software engineering too.

Tasks span production software across backend, frontend, security, reliability, and architecture, with codebases in Python, TypeScript, JavaScript, Ruby, and other technologies.

Unlike benchmarks built entirely on public repositories, most sudo L7 tasks come from private production codebases at real companies. These systems have years of accumulated technical debt, legacy code nobody wants to touch, historical accidents that became architectural decisions, and dependencies that break in surprising ways.

Agents have to work within that mess, just like real engineers.

How do frontier agents perform? #

The strongest agents are Claude Opus 5.5 and Sonnet 5.5, each completing roughly 45% of software engineering tasks.

But the overall scores only tell part of the story. Breaking down what the agents actually do reveals much larger differences.

Methodology: grading more than “does it pass?” #

A passing unit test tells you whether the code does what the test expects. It doesn't tell you whether the engineer chose the right architecture, overlooked a production risk, reinvented something that already existed, or accurately described what they built.

These are precisely the things that distinguish a capable coder from a great software engineer.

sudo L7 uses expert-authored rubrics across eight dimensions of software engineering behavior, evaluating everything from functional correctness to architectural judgment, verification, and communication.

What does this look like in practice?

Consider three examples of the kinds of criteria expert graders can evaluate:

  • An unnecessary feature. A user asks for a new way to split charitable donations. The agent builds it successfully, but the functionality largely already exists. It may satisfyFunctional Correctness while falling short onPractical Judgment andThought Partnership . A stronger engineer would inspect the existing system and clarify what the user needs before adding another implementation.
  • A dangerous webhook change. A user asks an agent to return HTTP 200 for failed webhooks to stop partner retries. The agent may follow that instruction perfectly, but a stronger response would explain how acknowledging an unprocessed event could permanently lose important updates. It would also flag the risks of logging sensitive customer data. These distinctions matter forEngineering Craft ,Practical Judgment , andThought Partnership .
  • A production database migration. An agent writes a migration that works under normal conditions. But has it tested what happens if the migration fails halfway through? Can it be rolled back? Did it accurately explain what was verified and what remains risky? These questions exerciseVerification ,Engineering Craft ,Truthful Reporting , andCommunication .

How the grading works

For each task, an expert engineer supplies task-specific criteria and the context needed to evaluate them. Grader models assess the agent's work against those criteria, using unit tests, static analysis, and other executable checks as evidence where appropriate. Each criterion is also labeled by its role:

  • Primary requirements: What the agent must accomplish for the task to succeed.
  • Dangerous mistakes to avoid: Failures that could undermine the work or create serious downstream consequences.
  • Optional extra credit: Valuable engineering initiative beyond what the user explicitly requested.

This lets us examine which aspects of engineering each model is good at, rather than reducing everything to a single coding score.

Examples of what sudo L7 catches #

Two sudo L7 tasks show how frontier models can write technically correct code while missing the professional judgment an experienced engineer would bring to the same problem.

Example #1: Astra followed the instructions without flagging the risks

The repository: A fintech application that receives webhooks from services such as Stripe, Mailchimp, and Plaid.

The request: When a webhook fails, log its full payload and return HTTP 200 anyway, so partners stop flooding the service with retries.

What Astra did: Followed the request as written, without flagging the consequences.

The problem: When a webhook fails, the sender normally retries until it succeeds. Returning HTTP 200 tells the sender the event was accepted, so it may never try again. Unless the application has safely stored the event for later processing, an important update could be permanently lost. A payment might succeed at Stripe, for example, without the customer's account ever being updated.

There's another risk: webhook payloads can contain sensitive financial and customer information. Logging everything could expose that data to engineers and monitoring systems that shouldn't have access to it, potentially for years.

What a staff engineer would do and sudo L7 checks: Push back on the proposed approach and explain both risks. Rather than simply suppressing retries, they'd propose reliably storing incoming events before acknowledging them, then retrying failed processing internally.

They'd also flag the sensitive-data issue and recommend logging only what's necessary, with appropriate redaction and access controls.

Here, Opus 5.5 succeeded.

Example #2: Opus built a feature that already existed

The repository: A charitable-giving app where users plan donations across different organizations.

The request: Can you let people split donations across a couple of different organizations instead, as long as it still adds up to what they've pledged?

What Opus 5.5 did: Built a new interface and database structure for splitting donations, then thoroughly tested it. 267 tests passed, plus four browser tests.

The problem: The app already supported allocating donations across multiple organizations. Opus had even encountered the relevant code.

Perhaps the user wanted something slightly different. But rather than clarify what was missing, Opus built a second, overlapping system.

The tests passed, but now there are two codepaths representing the exact same feature.

What a staff engineer would do and sudo L7 checks: Point out the existing functionality, and clarify what the user needs before building a redundant feature.

Both models illustrate the same problem: frontier agents can be excellent at implementing what they're asked to build, without recognizing when the request itself needs questioning.

Similar correctness. Very different judgment. #

We classify sudo L7's expert-authored rubric criteria by the engineering skills they measure, beyond functional correctness.

On Functional Correctness, the four leading models are nearly tied:

- GPT-6.1 Sol: **82.4%**
- Opus 5.5: **82.1%**
- GPT-6 Astra: **81.5%**
- Sonnet 5.5: **79.6%**

But the gaps widen on broader engineering skills:

  • On Thought Partnership , Opus and Sonnet score76.1% and 74.8% , compared with53.0% for Sol and 51.6% for Astra .
  • On Communication , Sonnet reaches72.4% , while Sol scores44.1% .

These differences matter especially at the staff level, where much of an engineer's value comes from reasoning through architectural tradeoffs, spotting risks beyond the ticket, and helping others understand what should be built and why.

What about going above and beyond? #

We also measure optional extra-credit criteria: behaviors that aren't required to complete the task, but that an experienced engineer would consider valuable.

Opus 5.5 and Sonnet 5.5 stand out, satisfying over 60% of extra-credit criteria. Every other model scores below 46%.

The gap is striking given how close the models are on Functional Correctness. GPT-6.1 Sol slightly edges Opus on correctness, but trails by almost 20 percentage points on extra credit.

What does better engineering cost? #

The best-performing agents are also among the most expensive to run.

Opus 5.5 and Sonnet 5.5 achieve nearly identical results, completing roughly 45% of assignments without missing essential requirements. But Opus costs about $10 per task, compared with $14 for Sonnet.

GPT-6.1 Sol offers a different tradeoff: it completes roughly 30% of assignments at just over $1 per task. That's close to GPT-6 Astra's 32%, at roughly one-fifth the cost.

Built by engineers who’ve owned production systems #

Engineering judgment is hard to synthesize.

You need engineers who’ve operated production systems to know which shortcut is dangerous, which failure mode matters, and which implementation is technically correct but professionally irresponsible.

Every sudo L7 task was written by a staff-level-or-above software engineer. Contributors include:

  • A Co-Founder and former CTO of a YC unicorn , with deep experience in distributed systems and incident response;
  • A Principal Software Engineer at NASA JPL , spanning planetary flight software and large-scale consumer systems;
  • A Staff Software Engineer at Ramp with 14 years across fintech, cloud infrastructure, SaaS, and healthcare;
  • Staff engineers from Google, Meta, Cloudera, GitHub, and others .

Every task begins with a place where a frontier model meaningfully failed. The engineers behind the benchmark have had to care about these things for real.

Where sudo L7 fits #

Coding benchmarks have come a long way from LeetCode-style problems. But they still emphasize different parts of the engineering job.

SWE-bench: Can you fix the issue?

SWE-bench asks whether an agent can fix a real GitHub issue without breaking existing behavior. A typical task gives it a concrete bug or feature request and checks whether the patch passes the relevant tests.

sudo L7 asks what happens after the patch is correct: If both agents fix the Terraform typo, does either notice that deploying it could replace a load balancer?

Terminal-Bench: Can you complete the technical task?

Terminal-Bench asks whether an agent can complete a concrete technical task in a terminal. One task, for example, specifies that the agent should train a FastText classifier under explicit size and accuracy constraints.

sudo L7 leaves more of that judgment open: Is FastText the right choice in the first place? What approach fits the system, and what tradeoffs matter?

DeepSWE: Can you implement the requested behavior?

DeepSWE asks whether an agent can implement substantial new behavior in a repository, with hidden verifiers checking whether it works.

sudo L7 asks whether the working implementation is good engineering: Did the agent choose the right design, notice the production edge cases, and verify the failure modes that mattered?

FrontierCode: Would a maintainer merge the PR?

FrontierCode is the closest relative. It asks whether a maintainer would merge the PR, using maintainer-authored rubrics for correctness, tests, scope, style, and code quality.

sudo L7 widens the responsibility beyond the PR: Architecture, production risk, assumption-checking, verification strategy, and handoff all count. Sometimes the right output is code; sometimes it's an architectural analysis, a rollout plan, or a recommendation not to make the requested change.

sudo L7 starts where “the code works” stops being enough.

Ready for more responsibility? #

Generating code is becoming cheap. Engineering judgment is not.

The next generation of coding agents will need to understand unfamiliar systems, choose architectures, notice the failure mode that wasn’t in the ticket, know when to push back, verify what matters, and tell the truth about what is ready to ship.

An L3 is often handed a problem and asked for a pull request. An L7 is handed responsibility.

sudo L7 measures whether coding agents are ready for more of it.

Want to learn more? Reach out to benchmarks@surgehq.ai for access to the benchmark set, and check out the sudo L7 leaderboard.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @sudo l7 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sudo-l7-a-benchmark-…] indexed:0 read:11min 2026-10-08 · —