# Accelerating agentic RL and evaluation research velocity with 45x faster GKE Agent Sandbox

> Source: <https://cloud.google.com/blog/products/containers-kubernetes/accelerate-agentic-rl-with-gke-agent-sandbox/>
> Published: 2026-09-29 17:00:00+00:00

When scaling up agentic reinforcement learning (RL) and evaluation across massive parallel rollouts, frontier AI labs inevitably hit a bottleneck: Expensive GPU clusters sit idle, waiting minutes for CPU sandbox cold-starts, plus thousands of multi-gigabyte [SWE-bench](https://www.swebench.com/original.html)-style image pulls and scheduling backlogs. It’s a sandbox infrastructure problem that silently slows down your research and burns your training budget.

To solve this fundamental infrastructure bottleneck, **today we are introducing** **GKE Agent Sandbox****optimized for RL** along with the [**Agent Sandbox RL orchestration SDK**](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/examples/agent-sandbox-rl), plus native [integrations](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/clients/integrations) for popular RL gyms and harnesses, now generally available.

As the operating system for modern AI, Kubernetes has evolved to power massive GPU/TPU training clusters and distributed inference. Now Kubernetes is expanding to drive the next AI compute frontier: agents. But unlike static workloads, agentic workloads evolve rapidly, so infrastructure must evolve just as fast. Rather than guessing at what RL researchers needed, **we placed Kubernetes itself on an** **auto-research and verification loop driven by performance benchmarks and evaluations**. We used heavy agentic benchmarks like [SWE-bench](https://www.swebench.com/verified.html) to intentionally stress-test and break our own clusters. Every bottleneck that surfaced — from etcd timeouts to GPU idle spikes — was fed back into our development cycle to refine GKE’s core primitives.

This resulted in a purpose-built sandbox layer for agentic RL and eval workloads that features:

With this new primitive, AI labs and agent-native startups can now reliably run large scale agentic RL trajectories and evals simultaneously, minimizing accelerator idle time and drastically accelerating their research velocity.

Before we talk about the solution, let’s be precise about what makes agentic RL so demanding for infrastructure in the first place. In a standard agentic RL loop, an LLM policy generates actions like code snippets on GPUs and executes them inside isolated CPU sandboxes to observe a reward signal. However, when scaling up this loop to support tens of thousands of parallel rollouts, three critical infrastructure bottlenecks emerge:

These are not hypothetical problems. They are the daily reality for frontier AI labs training state-of-the-art agents.

To fan-out the code execution sandboxes, we set up a relatively modest cluster — a 10-node gVisor sandbox pool with GKE image streaming enabled, the [Agent Sandbox](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/machine-learning/agent-sandbox) controller, and an in-cluster [SDK](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/examples/agent-sandbox-rl) driver to claim warm pods. We tested various strategies and setups including a large number of images and high cardinality.

[GKE Agent Sandbox](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/machine-learning/agent-sandbox) is the open Kubernetes primitive for secure agent execution, featuring built-in SandboxWarmPool capabilities that eliminate cold-start overhead by maintaining pre-initialized, healthy environments. By integrating SandboxWarmPool with [GKE Image Streaming](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/image-streaming), we effectively support RL and eval workloads that demand high image cardinality — even with thousands of large images (>1.2GB), while delivering the exceptionally low TTFC required for responsive agentic training. Its [snapshot, suspend and resume](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/agent-sandbox-pod-snapshots) capabilities support checkpointing for error recovery, and sandbox forking for parallel agent branching logic to explore multiple trails.

This GKE primitive is already powering the agentic RL training infrastructure at Mistral AI, a frontier AI lab:

“To push the boundaries of reinforcement learning, you need infrastructure that can instantly scale to handle unpredictable demand. By leveraging GKE's high-performance Agent Sandbox for RL, we can seamlessly orchestrate hundreds of thousands of secure environments across clusters and handle spikes of over 30,000 sandboxes on a single cluster. It provides the reliable foundation we need to accelerate our model training and iteration cycles.” - Jean-Malo Delignon, Research Engineer, Mistral AI

To help researchers harness this power without writing Kubernetes YAML or embedding custom daemons into container images, we built the [Agent Sandbox RL orchestration SDK](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/examples/agent-sandbox-rl). It exposes a clean, async Python API with pluggable warm-pooling strategies tailored to your specific evaluation or RL training pattern. We also built native [integrations](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/clients/integrations) for RL tools like [Gymnasium](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/clients/integrations/gymnasium), [NVIDIA NeMo Gym](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/clients/integrations/nemo-gym), and [OpenHands](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/clients/integrations/openhands), with more to come.

We ran this on a 10-node gVisor sandbox pool against two workloads: a 500-image [SWE-bench](https://www.swebench.com/verified.html) environment, and a harsher 4,578-image [R2E](https://huggingface.co/R2E-Gym/datasets) corpus that does not fit in local disk.

**1.TTFC and tail latency: 10x–45x faster**

| **Core metric** | **Why it matters** | **Raw K8s pod baseline** | **GKE Agent Sandbox SDK** | **Benchmark gains** | 
|---|---|---|---|---|
| **TTFC, average** | Determines GPU idle time | 44s – 85s (average) | 1.1s – 8.8s (average) | **10x faster** | 
| **Tail latency — max TTFC (worst case)** | The bottleneck for the batch | 7.5 mins (450 seconds) | Strictly under 10 seconds | **45x faster** | 

Concurrency ranged from 500 simultaneous tasks and sandboxes up to 18,312 (4,578 R2E images × 4 rollouts). The gain held across every setup.

**How we got here.** We instrumented the controller, the SDK, and the RL fleet, then ran hundreds of comparison runs against that fixed 10-node budget. Two changes produced nearly all of the gain:

**Image hydration moved off the critical path.** A high-cardinality corpus far exceeds local disk, so pulling cold images mid-training creates I/O contention and multi-minute tails. The SDK plans image placement across nodes, and [GKE Image Streaming](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/image-streaming) plus upfront warm-pooling handle the rest — hydration finishes before the rollout ever asks for it.

**The control plane is paced.** High-concurrency rollouts trigger API-server thundering herds. [Agent Sandbox Controller v1.0.0](https://github.com/kubernetes-sigs/agent-sandbox/releases/tag/v1.0.0) adds rate controls that bound how fast warm pools refill, keeping etcd and the API server stable through the burst.

The tuning knobs, the benchmark harness, and the load tests behind these numbers all ship in the [agent-sandbox-rl](https://github.com/kubernetes-sigs/agent-sandbox/tree/main/examples/agent-sandbox-rl) example, which emits human-readable and JSON reports so you can reproduce the comparison on your own cluster.

**2. 3x less control-plane and pod lifecycle churn during multi-trajectory rollouts**

| **Core metric** | **Why it matters** | **Raw K8s pod baseline** | **GKE Agent Sandbox SDK** | **Benchmark gains** | 
|---|---|---|---|---|
| **Control-plane churn** (4,578 images × 4 rollouts run) | Control-plane saturation | 18,312 | 5,869 | **3.1x less churn** | 

Raw Kubernetes recreates a pod for every trajectory step. At 18,312 tasks that inflates scheduling overhead, wastes disk I/O, and — in our large-scale tests — triggered unbounded garbage-collection loops. The SDK's in-place `recycle` strategy keeps the pod alive instead, running an in-pod `git reset` and repository checkout between rollout episodes. Pod creations drop 3.1x and the control plane stays flat through the burst.

Warming environments up front spends inexpensive CPU and background cluster time so that image streaming and readiness checks never land on the critical path. For an RL fleet that’s a straightforward trade: Accelerator idle time is the expensive resource, and this approach eliminates it.

One detail matters at training scale: The SDK counts and surfaces every failed sandbox as retriable rather than silently dropping it. An untracked drop is not just a lost rollout — it is reward bias.
