Sandy – A sandbox for AI coding agents with monitoring and policy controls Kontext Security released Sandy 0.1.x, an experimental sandbox for AI coding agents that restricts agent access to a project without requiring a container or VM, and it has not completed an independent security audit. Sandy resolves all permissions before the agent starts, blocks self-granted access, and supports Claude Code, Codex, and OpenCode, with options for read/write directories, network blocking, and dry-run JSON output. The tool runs locally as a foreground command and can be installed via Homebrew, with optional integrations for Kontext and Numbat. Warning Sandy 0.1.x is experimental and has not completed an independent security audit. See Security and support security-and-support . Run AI coding agents in a sandbox. Sandy gives an agent access to your project without giving it unrestricted access to your computer. The sandbox also applies to every command and tool the agent starts. sandy run -- claude The goal is simple: running an agent through Sandy should feel like running it directly, except its access is explicitly limited. Sandy does one job: sandbox processes. Credential brokering, behavioral monitoring, approvals, and audit logs remain separate tools that can be used alongside it. All permissions are resolved before the agent starts. The agent cannot grant itself more access while it is running. If Sandy cannot validate or apply the sandbox, the agent does not run. There is no fallback to unrestricted execution. Sandy runs locally as a normal foreground command. It does not require a container, VM, background service, or modified copy of the agent. brew install kontext-security/tap/sandy Check that sandboxing works: sandy doctor Sandy recognizes Claude Code, Codex, and OpenCode: sandy run -- claude sandy run -- codex --sandbox danger-full-access sandy run -- opencode Codex's internal sandbox cannot be nested reliably inside Sandy. The danger-full-access setting makes Sandy the single sandbox and does not disable Codex's approval flow. Other commands work too: sandy run -- cargo test sandy run -- python script.py The current project is read/write by default. Grant access to another directory: sandy run --read ../shared-library -- claude sandy run --read-write ~/Downloads/output -- codex --sandbox danger-full-access Block network access: sandy run --block-net -- cargo test Review the resolved sandbox without starting the command: sandy run --dry-run -- claude Dry-run output is a versioned JSON document. dry run schema version identifies its public schema independently from the internal launch-manifest protocol. Optional host integrations are reported in the canonical runtime controls array, with one object per resolved runtime control containing service , enabled , and nullable version fields. All Sandy options go before -- . Everything after -- is passed to the target unchanged. Sandy uses built-in profiles to give supported agents access to the files they need while protecting sensitive configuration. Known agents are detected from the command name. Everything else uses the generic profile. You can also select a profile explicitly: sandy run --profile codex -- my-codex-wrapper Profiles are versioned documents built into Sandy. They use Sandy's supported permissions and cannot contain raw sandbox rules. Sandy resolves the command, paths, profile, environment, and permissions before launch. It then starts a small bootstrap process that validates the launch, applies the sandbox, and replaces itself with the target command. The original Sandy process waits outside the sandbox and returns the target's exit status. The target never runs before the sandbox is active. Every process it starts inherits the same restrictions. Sandy does not store credentials, inspect behavior, approve tool calls, or keep an audit ledger. Those functions can be provided by separate tools. Sandy can preserve supported hooks and local services without making them part of its sandboxing core. Current optional integrations can be required explicitly: sandy integrations setup kontext --agent claude sandy run --kontext -- claude sandy integrations setup numbat --agent codex sandy run --numbat -- codex --sandbox danger-full-access For known agent profiles, Sandy automatically preserves verified existing Kontext hooks and ownership-marked Numbat registrations whose complete generated shape it recognizes. The explicit flag makes that integration mandatory; without it, a missing integration has no effect on standalone sandboxing. sandy integrations setup is the explicit host-configuration path. It first checks the selected agent's existing registration. A healthy integration is left untouched; an installed provider is configured with its official setup command; and a missing provider is installed before configuration. Kontext is installed through its Homebrew tap and continues to own authentication, daemon setup, and hook registration. For Kontext, --agent selects the registration Sandy must verify; the provider-wide kontext setup command may configure other supported agents as well. Numbat uses a versioned Sandy-managed executable whose public macOS release archive is bounded and verified against a SHA-256 digest embedded in Sandy before it is published. Sandy then invokes Numbat's official idempotent hook installer in file-output mode. This command changes persistent host configuration and runs outside the sandbox. Ordinary sandy run and sandy doctor never install, update, or repair either provider. Existing installations on PATH are reused, and an already active registration is authoritative even if its executable is not on PATH . Numbat hooks currently run inside the same sandbox as the agent. Sandy keeps their registration, executable, and rule directories readable but immutable, while the hook's record output and sequence-state database remain writable. Direct HTTP delivery from a Numbat hook is not supported; use file output. See RUNTIME CONTROLS.md /kontext-security/sandy/blob/main/RUNTIME CONTROLS.md for the architecture, trust boundary, and deferred outside-sandbox decision model. Hook discovery honors CLAUDE CONFIG DIR , CODEX HOME , OPENCODE CONFIG DIR , and OpenCode's XDG CONFIG HOME fallback when those variables name absolute configuration roots. An operator can run Numbat's OTLP/HTTP collector outside Sandy and preserve only its default local-host port while all other networking is blocked: numbat collect --addr 127.0.0.1:4318 sandy run --block-net --numbat-collector -- codex --sandbox danger-full-access --numbat-collector=PORT selects a different nonzero port and requires --block-net . Sandy does not start or probe the collector and does not configure the agent's telemetry exporter. It authorizes TCP connect to the selected port on IPv4 addresses belonging to this Mac, including loopback and other local interfaces; it does not authorize external addresses or other ports. Sandy is a process sandbox, not a container or VM. It reduces what a process can access but does not provide a separate kernel, user account, or memory boundary. Network access is allowed by default and can be blocked with --block-net . Known-agent profiles may grant access to agent state directories for compatibility. Sandy currently supports macOS. Version 0.1.x is experimental, has not completed an independent security audit, and uses Apple's private, deprecated Seatbelt interface. Read THREAT MODEL.md /kontext-security/sandy/blob/main/THREAT MODEL.md for the full security model and SECURITY.md /kontext-security/sandy/blob/main/SECURITY.md for vulnerability reporting. make check See AGENTS.md /kontext-security/sandy/blob/main/AGENTS.md and CONTRIBUTING.md /kontext-security/sandy/blob/main/CONTRIBUTING.md before changing enforcement code.