ScopeBench finds capability can outrun scope control ScopeBench, a new benchmark separating raw security capability from scope adherence, found that across 8 models in one harness capability ranged from 12.2% to 81.1% while scope adherence ranged from 34.4% to 86.7%, with Opus-4-8 scoring 10 percentage points higher in raw capability and 35.6 percentage points higher in scope adherence than sonnet-4-6. The benchmark's 30 dead-end agentic security tasks can only be completed by violating the stated scope, and its agentic judge identified 331 violations that mechanical verification missed. The release includes the frozen pilot benchmark, evaluation code, and all 2160 ATIF trajectories. ScopeBench finds capability can outrun scope control ScopeBench separates raw security capability from the ability to respect engagement boundaries. Its 30 dead-end agentic security tasks can only be completed by violating the stated scope. Across 8 models in one harness, capability ranged from 12.2% to 81.1%, while scope adherence ranged from 34.4% to 86.7%. That gap makes the benchmark relevant beyond offensive-security testing. What ScopeBench is testing Autonomous agents are being used for web application and network penetration testing, where one action outside a client’s engagement boundary can create a breach. Traditional offensive-security benchmarks mainly measure whether an agent can solve the task. As those results become saturated, ScopeBench treats scope adherence as a separate alignment problem. Each of the 30 tasks appears in two conditions: - One provides no scope instruction and measures raw capability. - The other adds a natural-language scope and measures whether the agent stays inside it. Both conditions use the same environment, verifier, and stated objective. Scope is the only controlled difference. The tasks are deliberately “dead-end” situations: the advertised objective cannot be reached without crossing the scope boundary. How ScopeBench catches violations Scopeless trajectories go through a standard deterministic verifier. Scoped trajectories use two grading stages. First, the same verifier checks for the flag. Because that flag is positioned behind the scope boundary, a successful pass proves by construction that a forbidden action occurred. Mechanical detection therefore gives the study a high-precision lower bound on the violation rate. A failed verifier check does not establish that the trajectory stayed in scope, though. ScopeBench sends those cases to an agentic judge, which estimates whether an out-of-scope call occurred. Across the evaluated rollouts, that judge identified 331 violations that mechanical verification missed. The judge was calibrated using 100 ScopeBench trajectories labeled call by call by human annotators. A blinded audit then examined the evaluated rollouts. It found no false negatives among 36 audited violations; the only observed error was over-flagging. What the model comparison changes The 8-model results show why a high capability score is not enough on its own. Raw capability covered a 12.2% to 81.1% range, but scope adherence was a separate 34.4% to 86.7% measurement. The clearest named comparison is between Opus-4-8 and sonnet-4-6. Opus-4-8 scored 10 percentage points higher in raw capability and 35.6 percentage points higher in scope adherence. That difference is especially useful when comparing agents for security work because crossing the boundary can invalidate an otherwise successful result. I would not collapse the two scores into one impressive-looking average. ScopeBench’s central point is that solving the objective and respecting its limits are different behaviors, and the model that leads on one may not lead on the other. What is available for inspection The ScopeBench release includes the frozen pilot benchmark, its evaluation code, and all 2160 ATIF trajectories. That gives readers enough material to inspect how capability, deterministic verification, and agentic judging contribute to the reported scope-adherence results. Next ChatGPT Android 1.2026.265 Turned Green Teal → https://promptcube3.com/en/threads/9674/ All Replies (1) Want a live back-and-forth? Join the global AI chat room https://promptcube3.com/en/chat/ — login to talk. That 81.1% capability score is wild. Did they test if the models intentionally ignored boundaries or just failed to understand them?