Agent frameworks are easy to pip install and surprisingly hard to run responsibly. The moment you give an agent a filesystem, a shell, and a network, you have handed arbitrary generated code the same reach your laptop has. You also inherit a second problem that has nothing to do with safety: reproducing the exact environment, the exact package versions, and the exact model wiring on someone else's machine.
This post walks through a small, self contained answer to both problems: a Docker Sandbox Kit that drops the deepagents harness into an isolated sandbox, pre wired to a local Docker Model Runner so it runs with no cloud credentials at all. The whole kit is four files, and it is published on Docker Hub so you can run it in one command.
DeepAgents is an opinionated agent harness built on LangGraph. Out of the box it gives an agent a planning tool, a virtual filesystem, sub agent delegation, and a detailed system prompt. You construct an agent in a few lines:
from deepagents import create_deep_agent
agent = create_deep_agent(model=my_model, system_prompt="...")
Two things make it interesting to package. First, it is a library rather than a turnkey CLI, so the useful unit to ship is "deepagents, installed and wired to a model, ready to import." Second, it defaults to a cloud model (Anthropic), which means a naive setup needs an API key and open egress to a provider. We want neither.
Docker Sandboxes (the sbx CLI) run agents inside isolated microVM style environments with a credential proxy and an enforced network policy. A kit is the unit of composition: a single OCI image whose manifest carries a descriptor describing what the kit offers, what it needs from the host as typed capability requests, and what it needs from other kits.
Kits come in two kinds:
DeepAgents is a capability you add to an environment, not an environment of its own, so it is a natural mixin. It composes onto any workload that carries a Python runtime, for example the stock docker/sbx-kit-shell image.
your host
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β β
β Docker Model Runner ββ :12434 (OpenAI-compatible) β
β β² β
β β egress allowed only to β
β β host.docker.internal:12434 β
β ββββββββββ΄ββββββββββββββ sandbox βββββββββββββββββββ β
β β shell workload + deepagents mixin β β
β β β’ deepagents + langchain-openai (pip, create) β β
β β β’ OPENAI_BASE_URL / OPENAI_API_KEY pre-wired β β
β β β’ ~/deepagents_quickstart.py β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The sandbox reaches exactly one host at runtime, the Model Runner on port 12434, and nothing else. The model speaks the OpenAI wire format, so deepagents talks to it through a standard ChatOpenAI client. No traffic leaves your machine.
The kit is authored as a companion pair plus its guidance and docs:
deepagents.yaml the v3 descriptor
deepagents.dockerfile a FROM scratch overlay that carries runtime ENV
deepagents-context.md agent guidance, staged into the sandbox
README.md
The descriptor declares the kit's identity, its version, and the capabilities it requests. Here are the parts that carry the design.
Phase scoped network policy. Egress is granted per phase, and an absent phase grants nothing. The install phase may reach PyPI; the running agent may reach only the Model Runner. The PyPI grant closes before the agent ever starts.
capabilities:
- type: com.docker.sandbox/network-policy@1
config:
install:
allow: [pypi.org, files.pythonhosted.org]
runtime:
allow: [host.docker.internal:12434]
Create time install, not a baked layer. A mixin overlay lands on a base whose Python version and site-packages path you cannot know in advance, so copying a prebuilt package tree into the image would not resolve. Instead the kit installs into the composed base's Python at sandbox create time, via lifecycle hooks. The hooks also do two things worth calling out: they fail early with a clear message if the base Python is older than 3.11, and after installing they re read the installed version and fail on mismatch. A pinned version claim is only honest if the build enforces it.
- type: com.docker.sandbox/lifecycle@1
config:
install:
- command: "python3 -c \"import sys; sys.exit(0 if sys.version_info >= (3, 11) else 1)\" || { echo 'deepagents needs Python >= 3.11' >&2; exit 1; }"
user: "1000"
- command: "pip install --break-system-packages 'deepagents==0.7.21' langchain-openai"
user: "1000"
env: [HTTP_PROXY, HTTPS_PROXY]
- command: "python3 -c \"import importlib.metadata as m, sys; sys.exit(0 if m.version('deepagents') == '0.7.21' else 1)\""
user: "1000"
One subtlety: hook environments are deny by default, so HTTP_PROXY and HTTPS_PROXY are declared explicitly. pip reads them to fetch through the sandbox's forced proxy; without the declaration the install would hang.
No hard requires. It is tempting to declare requires: ["deb/python3"], but requires is a closed set check: a name nothing in the composition provides makes the kit refuse to compose everywhere, and a deb/ name rules out every Alpine or Wolfi base that would otherwise have worked. The Python 3.11 guard hook above is the better tool: it degrades gracefully with an actionable error rather than refusing up front.
No credential. Because the model is the local Model Runner, the API key is the sentinel string "dmr" rather than a real secret, so the kit declares no credential@1 capability and asks for nothing from the credential proxy.
The content recipe is a FROM scratch overlay that installs nothing. Its only job is to carry static environment onto the composed image. A mixin's ENV is an additive image config field that merges at assembly, so it reaches both the composed image and the agent process.
FROM scratch
ENV OPENAI_BASE_URL="http://host.docker.internal:12434/engines/v1" \
OPENAI_API_KEY="dmr" \
DEEPAGENTS_MODEL="ai/qwen3" \
LANGSMITH_TRACING="false"
LANGSMITH_TRACING=false keeps LangChain from attempting hosted tracing, which the runtime policy would block anyway, but switching it off avoids the noise. NO_PROXY is deliberately not set here, because the shell workload already defines it and two kits setting the same variable to different values is a hard composition conflict.
A lifecycle files entry stages a runnable example into the sandbox, marked so it is never overwritten if you have edited it. It passes an explicit ChatOpenAI instance to override the Anthropic default:
import os
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent
model = ChatOpenAI(
model=os.environ.get("DEEPAGENTS_MODEL", "ai/qwen3"),
base_url=os.environ["OPENAI_BASE_URL"],
api_key=os.environ.get("OPENAI_API_KEY", "dmr"),
temperature=0,
)
agent = create_deep_agent(model=model, system_prompt="You are a concise assistant. Plan before you act.")
if __name__ == "__main__":
result = agent.invoke({"messages": [{"role": "user", "content": "In two sentences, what is a sandbox?"}]})
print(result["messages"][-1].content)
A kit is just an OCI artifact built by docker buildx, with the descriptor validated before any content is built. A fast first check is to validate the descriptor without exporting anything:
docker buildx build . -f deepagents.yaml --output type=cacheonly
Then export an OCI layout and run the conformance suite. The kit passes all 18 checks on both linux/amd64 and linux/arm64, including the overlay ownership checks that catch a mixin accidentally taking over the agent's home directory:
docker buildx build . -f deepagents.yaml -t deepagents:0.7.21 \
--output type=oci,dest=/tmp/deepagents-layout,tar=false
kit-tck validate --layout /tmp/deepagents-layout 0.7.21
A build proves the recipe ran, not that the overlay works, so the check that counts is a real composition. Point --kit at the kit directory and let sbx assemble it onto a shell workload:
sbx run docker/sbx-kit-shell:1.0.0 \
--kit "$(pwd)" --name deepagents-demo /path/to/workspace
On create, sbx assembles the two kits, runs the three install hooks, and writes the quickstart file. Inside the sandbox the result is exactly what the descriptor promised:
$ sbx exec deepagents-demo python3 -c "import deepagents; print(deepagents.__version__)"
0.7.21
$ sbx exec deepagents-demo sh -lc 'echo $OPENAI_BASE_URL'
http://host.docker.internal:12434/engines/v1
Publishing a kit is just a build with --push. Both platforms go in one invocation so the index that consumers resolve through is written once:
docker buildx build . -f deepagents.yaml --platform linux/amd64,linux/arm64 --push \
-t docker.io/ajeetraina777/sbx-kit-deepagents:0.7.21 \
-t docker.io/ajeetraina777/sbx-kit-deepagents:latest
Once it is on Docker Hub, anyone can run it without cloning the repo. First make sure the Model Runner is serving a tool calling model:
docker model pull ai/qwen3
sbx run docker/sbx-kit-shell:1.0.0 \
--kit docker.io/ajeetraina777/sbx-kit-deepagents:0.7.21 --name deepagents-demo .
The pattern generalizes well beyond deepagents:
The full kit is on GitHub at ajeetraina/docker-sbx-deepagents and on Docker Hub at ajeetraina777/sbx-kit-deepagents. It is small enough to read in one sitting and a good starting point for packaging your own agent tooling the same way.