Your agent hits 50% on SWE-bench. You celebrate, then face the real problem: now what? Training harder on the same tasks just accelerates overfitting. Building new benchmarks from scratch costs weeks of engineering time you do not have. Most agent teams stall here and call it a capability ceiling. Google Research open-sourced EnvHarness on October 2, and it is a fundamentally different answer: instead of new benchmarks, reshape the existing ones around your agent’s specific weaknesses.
The Core Idea: Adapt the Environment, Not the Task #
EnvHarness is a programmable wrapper layer. It sits between your agent and a training benchmark and changes three things: where the agent starts, what it is allowed to do, and how tasks are sequenced. The benchmark’s verifier — the code that actually checks whether a solution is correct — stays completely untouched. That is the key design decision. Verifiers like SWE-bench‘s are years old, battle-tested, and trusted by the community. EnvHarness does not touch them. It reshapes the context around the task, not the task itself.
The formulation is direct: adapted environment = static environment + EnvHarness. Wrap your existing benchmark, dial up difficulty where your agent is weak, and the same verifier still grades the result.
Stage, Contract, Chain: Three Controls #
EnvHarness exposes three modification types. Each targets a different aspect of the training environment.
Stage repositions initial conditions. For a household task where the agent must put a mug on a table: baseline has the mug already visible. With Stage, it is inside a closed drawer. Same goal, harder start, same verifier. The agent must now learn to search before reaching for an object.
Contract filters available actions and controls what the agent can observe. The most useful example for software agents: an agent that consistently submits patches without running tests. A Contract intercepts the submit action and requires test execution first. The agent does not know the rules changed. The verifier does not know either. You just closed a training shortcut the agent had learned to exploit.
Chain sequences multiple tasks with a shared action budget. Instead of one isolated problem, the agent gets task A then task B with context carried forward. This directly tests something static benchmarks cannot: whether an agent can maintain goals across longer trajectories without losing track of what it was doing three steps ago.
EnvRigger: The Automated Diagnosis Loop #
Designing Stage, Contract, and Chain modifications manually would make EnvHarness a research toy. EnvRigger is what makes it a developer tool. It runs your agent against the current environment, identifies recurring failure patterns, proposes targeted modifications, validates them across up to five test runs, and outputs a deployable modification — no human required in the diagnosis step. You review the output and decide whether to apply it. That is the right division of labor for production teams.
What the Benchmarks Actually Show #
The paper tests EnvHarness across five benchmarks in four domains: ALFWorld, WebArena, SWE-bench Verified, OfficeQA, and SpreadsheetBench.
- SWE-bench Verified: 54.79% versus 47.67% baseline — a 7.12 percentage point gain
- ALFWorld: 68.3% versus 62.4% — 5.9 points better
- Out-of-distribution ALFWorld tasks: +9.0 points (70.4% versus 61.4%)
- Average SWE-bench trajectory: 55.01 steps reduced to 49.61 — 9.8% more efficient
The out-of-distribution number is the one worth watching. A 9-point gain on tasks the agent was never trained on suggests genuine capability improvement rather than benchmark gaming. GenEnv, which generates entirely new training environments from scratch, scored lower on ALFWorld. SWE-smith, which synthesizes new SWE-bench problems, placed below EnvHarness while requiring more compute. EnvHarness beat both while keeping the original verifiers intact.
Getting Started #
The repository is at github.com/google-research/envharness under Apache 2.0. For SWE-bench, ALFWorld, and WebArena, pre-built Bridge implementations are included. To integrate a custom environment, you implement the ActionableEnv interface:
class MyEnvBridge(ActionableEnv):
def reset(self) -> Observation:
def step(self, action: Action) -> Tuple[Observation, float, bool]:
For containerized setups, the Bridge attaches externally above Docker or Kubernetes — no modification to the environment container. The repository also includes the full reinforcement learning implementation from the paper, which is useful for teams replicating the training setup rather than just the environment wrapper.
Where It Does Not Work #
Two hard limits. EnvHarness requires fast, cheap state restoration — production databases, physical robots, and infrastructure with real-world side effects are out. EnvRigger also runs multiple inference passes per diagnosed failure mode, so for expensive frontier models, budget 3 to 5 agent runs per modification cycle. Use it on the training evaluation loop, not production.
The Bigger Picture #
The agent training ecosystem has a structural problem: static benchmarks get saturated. SWE-bench Verified now sits above 50% for multiple frontier models. The default response is to keep making benchmarks harder, but that is an arms race that moves the goalposts rather than builds better agents. EnvHarness proposes a different model: the environment adapts to the agent’s weaknesses. It is Apache 2.0 and integrates with benchmarks teams are already using — no switching cost, no new toolchain. For teams hitting a ceiling on agent training, the paper is worth reading and the code is worth running this week.