OpenAI just released Codex Security as an open-source TypeScript SDK with 11K+ stars. It's not another static analysis tool. It's an agent-driven security scanner that orchestrates parallel discovery workers, validates findings with LLM reasoning, generates patches, and verifies fixes autonomously. This is what happens when security tooling becomes agentic rather than rule-based.
Traditional SAST tools run pattern matchers and dump alerts. Codex Security runs a multi-stage pipeline: discovery agents scan code, validation agents filter false positives, patch agents generate fixes, and verification agents confirm the patches work. Each stage has its own state management, tool boundaries, and failure modes. Here's how the plumbing works.
Codex Security splits discovery into worker pools. When you run cs scan . on a repository, the CLI spawns multiple discovery workers that operate concurrently across different code paths. Each worker is a Node.js process running Python analysis tools underneath.
The orchestration layer manages:
~/.codex-security/scans/.
The CLI requires Node.js 22.13.0+ (within 22.x) or 24.x/26.x, plus Python 3.10+. Python 3.10 also needs tomli for TOML parsing. This dual-runtime requirement exists because the discovery layer uses Python-based security tools (likely Semgrep, Bandit, or similar) while the orchestration and LLM integration run in TypeScript.
// Simplified worker coordination pattern
interface ScanWorker {
id: string;
targetPaths: string[];
findings: CandidateFinding[];
}
async function runDeepScan(repoPath: string): Promise<Finding[]> {
const workers = allocateWorkers(repoPath);
const results = await Promise.all(
workers.map(w => runWorker(w.targetPaths))
);
const merged = deduplicateFindings(results.flat());
return validateFindings(merged);
}
Discovery produces candidate findings. Validation decides which candidates are real vulnerabilities. This is where the agent layer kicks in.
The validation agent:
SECURITY.md policy from the repository.
Tool boundaries matter here. The validation agent does not have write access to the repository. It can read code and policy files, but it cannot modify source or commit patches. This separation prevents a validation bug from corrupting the codebase.
The CLI supports custom severity rubrics. You can define your own risk scoring logic and pass it to the validation stage. This lets teams align agent decisions with internal security standards instead of relying on generic CVE scores.
Once a finding is validated, the patch agent generates a fix. This is a separate agent with different tool access:
The patch generation flow:
Verification is critical. The agent can run unit tests, linters, or type checkers to validate the patch. If verification fails, the agent can retry with a modified patch or flag the finding as requiring manual intervention.
Failure modes at this stage:
Codex Security integrates into CI pipelines by running in headless mode. Set OPENAI_API_KEY or CODEX_API_KEY in the environment, then run cs scan . in your build container.
For remote or headless machines, use cs login --device-auth if your workspace allows it, or sign in over SSH with port forwarding. The CLI supports SSH agent forwarding to authenticate without storing credentials on the CI runner.
Export formats:
Observability hooks let teams audit agent decisions:
This is useful when debugging why a finding was validated or rejected. You can replay the validation step with different context or rubrics.
Codex Security stores all scan state locally in ~/.codex-security/scans/. Each scan session includes:
This local-first design means you can browse past scans, compare findings across versions, and re-run validation without re-scanning. The CLI provides cs browse to navigate saved sessions.
State isolation is important. Each scan session is independent. If you run multiple scans concurrently (e.g., on different branches), they don't interfere with each other. The session ID is derived from the scan start time and repository path.
The agent architecture enforces strict boundaries:
| Agent Stage | Read Access | Write Access | Network Access |
|---|---|---|---|
| Discovery | Repository files | Scan session only | None |
| Validation | Repository + policy | Scan session only | OpenAI API |
| Patch Generation | Repository + Git history | Temp branch only | OpenAI API |
| Verification | Temp branch | Scan session only | None (isolated tests) |
Discovery and verification agents do not call external APIs. Only validation and patch generation agents communicate with OpenAI. This limits the attack surface if an agent is compromised.
The CLI also supports Daybreak Blue access for customers with access to OpenAI's advanced cybersecurity program. Use --cyber-access-program daybreak_blue to enable it. Otherwise, omit the flag or use --cyber-access-program standard.
Codex Security supports three deployment shapes:
npm install --global @openai/codex-security, run cs scan . from your repo.npx @openai/codex-security scan . in a Docker container with OPENAI_API_KEY set. Export SARIF and upload to GitHub Code Scanning.cs login --device-auth on a headless server, then schedule scans with cron or a workflow orchestrator.
For CI, the CLI detects when it's running in a non-interactive environment and skips prompts. It exits with a non-zero code if high-severity findings are detected, which fails the build.
The TypeScript SDK is also available as @openai/codex-security on npm. You can embed the scanning logic into custom tooling or dashboards. The SDK exposes the same orchestration primitives: runScan(), validateFindings(), generatePatch(), verifyPatch().
Agent-driven security scanning introduces new failure modes that static tools don't have:
The CLI includes duplicate detection to avoid re-validating the same finding across scans. This reduces API costs and speeds up incremental scans.
Use Codex Security when you need agent-driven validation and patching on top of traditional security scanning. It's a good fit for teams that:
Avoid it if:
The orchestration plumbing is solid. Parallel workers, isolated state, and strict tool boundaries make it production-ready. The main trade-off is LLM dependency: you're exchanging determinism for smarter validation and automated patching.