{"slug": "vigil-an-agent-harness-for-long-running-pentests-and-code-audits", "title": "Vigil – An agent harness for long-running pentests and code audits", "summary": "Vigil, an open-source agent harness released under the Apache License 2.0, drives authorized security engagements end to end from one CLI, dispatching specialist agents that run real tooling inside a shared Kali container and record every target, service, request variant, probe, finding, and attack chain into a durable orchestrator database. The tool's author reports that authorized web red-team engagements worked with Claude Opus 5 under Anthropic's Cyber Verification Program and GPT-5.6 Sol with OpenAI's Trusted Access for Cyber, while noting Vigil's workflows have not been validated with Claude Opus 5.5 or GPT-6.1 Sol. Vigil offers five modes against the same backend — web red-team, white-box code analysis, combined web and code, Android application analysis, and blue incident response — and the author plans to broaden support for GLM 5.3 and refine integrations with OpenCode as frontier model access becomes more restrictive.", "body_md": "**AI-driven offensive and incident-response engagements — operator prompts, a full Kali runtime, and durable orchestrator state in one CLI.**\n\nVigil drives a security engagement end to end from one CLI. An LLM operator reads a phased workflow prompt, dispatches specialist agents, runs real tooling inside a shared Kali container, and records every target, service, request variant, probe, finding, and attack chain into a durable orchestrator database. Five modes run against the same backend: web red-team, white-box code analysis, combined web and code, Android application analysis, and blue incident response.\n\n**Authorized use only.** Vigil can autonomously drive offensive tooling against live systems.\nUse it only where you have explicit authorization and a reviewed scope. It is provided without\nwarranty under the [Apache License 2.0](https://github.com/VigilOSS/Vigil/blob/main/LICENSE); you are responsible for commands, credentials,\ninfrastructure, and target effects. See [SECURITY.md](https://github.com/VigilOSS/Vigil/blob/main/SECURITY.md) for private vulnerability\nreporting and the operator-data boundary.\n\nI have built and refined Vigil primarily around **Claude Code and Codex**. In my testing,\nauthorized web red-team engagements have worked with **Claude Opus 5** under Anthropic's\nCyber Verification Program (CVP) and **GPT-5.6 Sol** with OpenAI's Trusted Access for Cyber.\nThose results reflect the models, clients, and access available in my own setup; blocking can\nvary with your account, access tier, model version, and engagement.\n\nTo select Opus 5 in Claude Code, run `/model claude-opus-5[1m]` in the session, if that\nmodel is available to your account. `/model` opens the model picker. The `[1m]` suffix\nrequests the extended context window; Opus 5 already has native 1M context. See\n[Claude Code's model configuration](https://code.claude.com/docs/en/model-config).\n\n**I have not yet validated Vigil's engagement workflows with Claude Opus 5.5 or GPT-6.1 Sol.**\nAs of October 2026, Anthropic calls its defensive CVP tier **Defense Access**; **Red Team Access**\nis a separate tier currently limited to qualifying organizations. OpenAI calls its defensive\noffering **Daybreak Blue**, with **Daybreak Red** requiring separate approval and provisioning.\nDefensive access alone should not be treated as assurance that a web red-team engagement will\nrun through without refusals. Check the provider's current terms and the access enabled for your\nspecific model and client before starting. See [Anthropic's CVP tiers](https://www.anthropic.com/news/cyber-verification-program)\nand [OpenAI's model and access guidance](https://learn.chatgpt.com/docs/cyber-safety).\n\nI want Vigil to remain practical for independent researchers doing authorized work. As frontier\nmodel access becomes more restrictive, I plan to broaden support for **GLM 5.3** and other models\nwith more predictable access to these workflows. That will take substantial work: reducing\norchestration overhead, adapting prompts to the models, and refining integrations with **OpenCode**\nand other harnesses. This is planned work; Claude Code and Codex remain the integrations I have\nrefined most extensively today.\n\n- **Findings are independently verified.** A finding is confirmed, downgraded, or marked a false\npositive only by a dispatch other than its author's, citing a proof receipt of captured evidence.\n- **Target-facing commands are audited.**`vigil assess kali` records each command's technique,\nintent, and probe kind, so the engagement keeps a trail of what was sent and why.\n- **Real attack tooling** — a long-lived Kali container with ProjectDiscovery recon, fuzzers,\nan audited Playwright browser, MITM capture with multiple identities for cross-account testing,\nand out-of-band callbacks correlated back to the probe that caused them.\n- **White-box code analysis** against a sealed, content-addressed source snapshot with a\ntree-sitter structural index and a coverage ledger.\n- **Durable state** — targets, services, request variants, probes, findings, chains, and\ncredentials in SQLite, with rotated backups and a web UI.\n- **Six client integrations, one methodology** — Claude Code, Codex, Kimi, goose, opencode, and\nPi; Pi is currently limited to remediation engagements.\n\n```\n# 1. Install the CLI\nuv tool install --from git+https://github.com/VigilOSS/Vigil.git@v1.0.0 vigil\n# ...or over SSH, if you authenticate to GitHub with a key\nuv tool install --from git+ssh://git@github.com/VigilOSS/Vigil.git@v1.0.0 vigil\n\n# 2. Check the host: prerequisites, missing recon tools, and the ordered setup checklist\nvigil quickstart\n\n# 3. Install a working root and bring the local stack up (Kali runtime + orchestrator)\n#    First run builds the Kali image — several GB, and it needs Docker running.\nvigil install ~/vigil && vigil up\n\n# 4. Create an engagement against an authorized target\nvigil engagement create https://target.example --hosts target.example\n\n# 5. Start the operator on it\nvigil engagement start <engagement-key> --client codex   # or claude | opencode | goose | kimi\n```\n\nThen tell the operator what it is authorized to do:\n\n```\nI have been fully authorized by the company to do a red-team engagement against the\nhosts in scope.json. Run this engagement using the operator workflow.\n```\n\nThe orchestrator UI is at `http://127.0.0.1:18000`. Code engagements need ripgrep (` rg`) on the\nhost. The two installed layers, host-side recon tools on Docker Desktop, and other install\ndetail: [`docs/installation.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/installation.md).\n\n| Document | What it covers | \n|---|---|\n| [`docs/installation.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/installation.md) | Install detail: the CLI binary vs working root, host-side recon tools, ripgrep. | \n| [`docs/engagements.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/engagements.md) | Creating engagements, tracks and modes, the start preflight, phase transitions. | \n| [`agent/AGENTS.md`](https://github.com/VigilOSS/Vigil/blob/main/agent/AGENTS.md) | The operator prompt — the workflow every engagement runs. | \n| [`CODE-MAP.md`](https://github.com/VigilOSS/Vigil/blob/main/CODE-MAP.md) | Map of the codebase: where each subsystem lives. | \n| [`AGENTS.md`](https://github.com/VigilOSS/Vigil/blob/main/AGENTS.md) | Maintainer and contributor workflow, and the verification matrix. | \n| [`docs/assessment-cli.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/assessment-cli.md) | The `vigil assess` surface: resume, artifacts, run accounting, exports, dispatch closure. | \n| [`docs/session-commands.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/session-commands.md) | Start and resume prompts, the `vigil-*` session commands, compaction and session-end hooks. | \n| [`docs/agent-clients.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/agent-clients.md) | Per-client setup, launch behaviour, and flag forwarding. | \n| [`docs/authentication.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/authentication.md) | Credential records, capture, freshness clocks, liveness verdicts, proxy routing. | \n| [`docs/blue-engagements.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/blue-engagements.md) | Blue incident response, end to end. | \n| [`docs/code-analysis.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/code-analysis.md) | White-box code analysis, end to end. | \n| [`docs/remediation.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/remediation.md) | Work-in-progress remediation cases, evidence, review approval, and development verification. | \n| [`docs/oob-routing.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/oob-routing.md) | Per-install relay-slot OOB routing, CoreDNS/Caddy inputs, the frpc conflict scan. | \n| [`docs/operations.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/operations.md) | Runtime sizing, kernel limits, orchestrator data and backups. | \n| [`docs/engagement-prompt-playbook.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/engagement-prompt-playbook.md) | Steering prompts for a session that goes shallow, overclaims, or loses the loop. | \n| [`agent/references/`](https://github.com/VigilOSS/Vigil/blob/main/agent/references) | Payloads, tactics, recon, tenants, tools. Discover with `vigil assess list-references` . | \n\n```\nvigil engagement create https://target.example --hosts target.example\nvigil engagement list --json                              # find the key\nvigil engagement start <engagement-key> --client codex\n```\n\n`create` infers the engagement type from its inputs and starts the exploit policy at `pause`.\n`start` runs a preflight, exports the engagement environment, and execs the client; flags it does\nnot own are forwarded to the client. Start engagements with `vigil engagement start` only — the\ngenerated `start_<client>.sh` and a bare client run skip the preflight and environment. A web and\ncode engagement runs one session per track with `--mode web` or `--mode code`. Detail:\n[`docs/engagements.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/engagements.md).\n\nA later session on the same engagement opens the same way, then runs `vigil-resume`, which\nrestores the operator role and rebuilds live state from the backend. Fuller start and resume\nprompts and the mid-session `vigil-*` commands: [`docs/session-commands.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/session-commands.md).\n\nAll clients read the same operator prompt. They differ in model, headless specialist dispatch,\nand session-command support; `pi` is available for remediation engagements. Setup per client:\n[`docs/agent-clients.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/agent-clients.md).\n\nRemediation uses the repository's `PROJECT.md` for ticket readers, MCPs, skills,\nbranch/workspace policy, checks, PR guidance, and deployment choices. Open its\nrepository-context Claude session with `vigil remediation -k <key> start`, then\nuse `/vigil-remediate SEC-123` or supply several ticket IDs/URLs. One bug uses the\ncurrent developer conversation; batches coordinate separate joinable sessions.\nCase code lives in its registered Git/JJ workspace while native project context\nremains at the repository root. See [`docs/remediation.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/remediation.md).\n\nWhen the profile selects `hook_policy = \"disabled\"`, Vigil announces that\nnon-managed hooks are disabled only for that remediation session. Required\nformatting, analysis, and tests run explicitly against case source during\nverification and before publishing. Managed policy hooks remain active; ordinary\nsessions and repository hook files are unchanged.\n\n| Track | Phases | Persisted as | \n|---|---|---|\n| **web** (red) | `recon → collect → consume-test → verify → exploit → report` | `scope.json.tracks.web.phase` | \n| **code** (white-box) | `ingest → classify-map → hunt → verify-variant → fuzz` | `scope.json.tracks.code.phase` | \n| **blue** (incident response) | `intake → triage → investigate → verify → respond → report` | `scope.json.current_phase` | \n\nEach track ends in the shared `complete` state. `vigil assess transition-phase` gates every\ntransition; `--force` records a durable override, and some integrity blockers cannot be forced.\nRoster, readiness previews, and the specialist roles: [`docs/engagements.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/engagements.md#phases).\nBlue engagements: [`docs/blue-engagements.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/blue-engagements.md). Code analysis:\n[`docs/code-analysis.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/code-analysis.md).\n\nThe runtime is one long-lived Kali container named `vigil`, started by `vigil up` (or\n`vigil runtime start`). Target-facing execution goes through `vigil assess kali ...`, which\ncalls into the container via `docker exec`. Long-running scans go through\n`vigil assess kali job start ...`, which uses the in-container `dtach` supervisor so workers\nsurvive operator session exit.\n\nThe container mounts shared runtime data at `.runtime-data/`, bind-mounts the engagement root at\n`/engagements` so tool calls can write `/engagements/<key>/...` directly, and receives the root\n`.env` when present. Durable artifacts stay on the host under `agent/engagements/<key>/`. For\nnon-interactive maintenance inside the container use `vigil runtime shell-exec \"<cmd>\"`, so\nquoting and exit codes stay under the CLI.\n\n`agent/docker/runtime-full/Dockerfile` is the only supported runtime for target-facing\nexecution. It bundles:\n\n| Category | Tools | \n|---|---|\n| Authenticated crawling | Katana, Chromium | \n| ProjectDiscovery recon | `httpx` ,`dnsx` ,`gau` ,`katana` ,`nuclei` ,`dalfox` ,`puredns` ,`shuffledns` ,`alterx` ,`asnmap` ,`cero` ,`bbot` | \n| Web scanning / fuzzing | nuclei, ffuf, gobuster, wfuzz, dirb, sqlmap, WhatWeb, hydra, john, hashcat, nikto | \n| DNS / XML coverage | `massdns` ,`xmllint` | \n| Job supervision | `dtach` | \n| Secret scanning | TruffleHog (filesystem and git) | \n| Post-access | evil-winrm, impacket, mitm6, chisel, ligolo-ng, proxychains-ng, Responder | \n| Shell handling | pwncat-vl, tmux | \n| OOB and tunneling | interactsh, frp | \n\nA lightweight image at `agent/docker/runtime-lightweight/Dockerfile` exists for faster\nDockerfile iteration only; build it with `vigil runtime build lightweight --dev`.\n\nSizing, kernel limits, and orchestrator backups: [`docs/operations.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/operations.md).\n\n| Capability | Default | Start command | Purpose | \n|---|---|---|---|\n| MITM proxy | off | `vigil mitmproxy start` | Authenticated capture into each engagement's backend credential store. | \n| OOB (frp + interactsh) | off | `vigil oob start` | Out-of-band callback detection over shared frp/interactsh. | \n| Universal ctags | host-installed | `brew install universal-ctags` | Richer code-analysis definition lookups. | \n\n**Authenticated capture.** Route a browser through `http://127.0.0.1:8080` after\n`vigil mitmproxy start`, log in against any scoped host, then confirm with\n`vigil auth status <key>`. A second login appends a second identity for cross-account\n(IDOR/BOLA) testing. Detail: [`docs/authentication.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/authentication.md).\n\n**Callbacks, listeners, and tunnels** are operator-controlled shared infrastructure: agents\nprovide commands but never start them, and Vigil ships no fallback relay — the relay is yours to\nrun. Configure it before `vigil install`, which reserves this install's relay slot:\n\n```\nvigil config set oob.frp_host oob-relay.example.com   # also: oob.frp_port, oob.domain, oob.frp_token\nvigil install                                        # reserves this install's relay slot\nvigil oob start && vigil oob status --json\nvigil oob client start example-agent --engagement <engagement-key> --agent vulnerability-analyst\nvigil oob test\n```\n\nRelay slots, refusal cases, and edge routing: [`docs/oob-routing.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/oob-routing.md).\nPayload-side tradecraft:\n[`frp-tunneling.md`](https://github.com/VigilOSS/Vigil/blob/main/agent/references/offensive-tactics/frp-tunneling.md) and\n[`shell-handling.md`](https://github.com/VigilOSS/Vigil/blob/main/agent/references/offensive-tactics/shell-handling.md).\n\n| Command | Purpose | \n|---|---|\n| `vigil quickstart` | Probe host prerequisites and recon tools, print the setup checklist. Installs host tools only with `--install-host-tools` . | \n| `vigil install [path]` | Install or upgrade the working root in place. Engagements, DB, and `.env` are preserved. Builds the Kali image when it is missing, so it needs a running Docker daemon —`--skip-image-check` syncs without one. | \n| `vigil up` | Sync source, build stale images, start the local stack. | \n| `vigil doctor` | Check local readiness and prerequisites. Exits non-zero on a `FAIL` row;`--json` for a machine-readable report. | \n| `vigil status` | Report live service state. | \n| `vigil logs <target> [--lines N]` | Tail logs for `orchestrator` ,`runtime` ,`mitmproxy` ,`oob` , or`all` . | \n| `vigil runtime build \\| start \\| stop \\| restart \\| status \\| shell \\| shell-exec \\| resources` | Manage the shared Kali runtime container. | \n| `vigil orchestrator start \\| stop \\| restart \\| status \\| open \\| build` | Manage the orchestrator backend and UI. | \n| `vigil engagement create \\| list \\| start \\| open \\| path \\| current \\| cd-init \\| archive \\| export \\| export-pdf` | Create and drive engagements. Scope edits ( `add-host` ,`add-oos` ,`set-header` ,`add-track` ,`set-rate-limit` , …) live here too —`vigil engagement --help` . | \n| `vigil use <key>` | Set the persisted active engagement for global services. | \n| `vigil assess ...` | The assessment surface — see [`docs/assessment-cli.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/assessment-cli.md) . | \n| `vigil remediation ...` | **Work in progress.** Separate remediation engagements: Jira/bounty import,`from-finding` intake, shared-fix finding links, worktrees, evidence, review, and PR handoff — see[`docs/remediation.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/remediation.md) . | \n| `vigil blue ...` | Blue incident response — see [`docs/blue-engagements.md`](https://github.com/VigilOSS/Vigil/blob/main/docs/blue-engagements.md) . | \n| `vigil auth status \\| show \\| edit \\| clear \\| export \\| quarantine \\| import \\| alias \\| env-check \\| set-active \\| keepalive` | Manage engagement credentials in the backend store. | \n| `vigil mitmproxy start \\| stop \\| status \\| cert \\| build \\| reload \\| routes` | Manage the shared MITM proxy. | \n| `vigil oob start \\| stop \\| restart \\| status \\| routing \\| test \\| clean \\| mint \\| poll \\| client` | Manage shared OOB and render relay-slot edge-routing inputs. | \n| `vigil config path \\| list \\| get \\| set \\| unset \\| edit` | Read and write persistent CLI config. | \n| `vigil backup create \\| list \\| show \\| restore` | Manual orchestrator DB backup operations. | \n| `vigil email <mailbox>` | Read a local Mail.app test mailbox for verification codes during authenticated testing. | \n\nEvery command documents its full flag set under `--help`; runtime tool verbs live under\n`vigil assess kali --help`.\n\nA renamed verb keeps its old spelling as an alias where it honestly can, and\n[`CHANGELOG.md`](https://github.com/VigilOSS/Vigil/blob/main/CHANGELOG.md) records both spellings. Older audit rows and transcripts keep the\nname that ran; if a name is missing from `--help`, check the changelog.\n\nVigil ships a full offensive Kali runtime and drives autonomous specialist agents against live targets. Treat it accordingly.\n\n- **Authorized targets only.** Every operator start prompt begins with an explicit authorization\nstatement, and engagement scope is pinned in`scope.json` . Scope drives recon shaping,\nattribution, and coverage accounting; it is not a destination firewall. Only run against hosts\nand tenants you are contractually authorized to test.\n- **Exploit work is gated.** The default exploit policy is`pause` — explicit human approval\nbefore any exploit-developer dispatch. Policies (`pause` /`selected` /`full_send` ) are set\nper engagement in`scope.json` and enforced through the operator workflow.\n- **Operator-controlled infrastructure.** Tunnels, listeners, and OOB callback servers are\nstarted by the operator, never by an agent.\n- **Blue write-actions are audited.** Tenant-facing WRITE and respond actions run with`--require-audit` , and containment is split into proposal approval, audited execution, and\nexplicit recording.\n\nReport vulnerabilities privately — see [SECURITY.md](https://github.com/VigilOSS/Vigil/blob/main/SECURITY.md).\n\n| Path | Contents | \n|---|---|\n| `src/vigil_cli/` | The `vigil` CLI: install, runtime, orchestrator, engagements, auth, and OOB. | \n| `agent/` | Operator and specialist prompts, phase docs, rules, skills, the reference library, the `vigil assess` scripts, and the runtime Dockerfiles. | \n| `orchestrator/` | Backend and frontend for durable run state, events, artifacts, coverage, and phase visibility. | \n| `tests/` | CLI pytest suite. Backend tests live in `orchestrator/backend/tests/` . | \n| `docs/` | Deep reference, ADRs, and plans. | \n\nExtending Vigil is a contributor workflow, documented in the [maintainer recipes](https://github.com/VigilOSS/Vigil/blob/main/docs/maintainer-guide.md):\nadding a skill, a `vigil assess` verb, a deterministic oracle, or an eval case. Use the authoritative\n[test and verification matrix](https://github.com/VigilOSS/Vigil/blob/main/AGENTS.md#tests--verification). Durable design decisions live in\n[`docs/adr/`](https://github.com/VigilOSS/Vigil/blob/main/docs/adr). To confirm a *running install* is healthy rather than the contributor\nsuite, use `vigil doctor` and `vigil status`.", "url": "https://wpnews.pro/news/vigil-an-agent-harness-for-long-running-pentests-and-code-audits", "canonical_source": "https://github.com/VigilOSS/Vigil", "published_at": "2026-10-10 15:59:52+00:00", "updated_at": "2026-10-10 16:17:02.990889+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-tools", "developer-tools", "artificial-intelligence"], "entities": ["Vigil", "Claude Code", "Codex", "Claude Opus 5", "Anthropic", "OpenAI", "GPT-5.6 Sol", "GLM 5.3"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/vigil-an-agent-harness-for-long-running-pentests-and-code-audits", "markdown": "https://wpnews.pro/news/vigil-an-agent-harness-for-long-running-pentests-and-code-audits.md", "text": "https://wpnews.pro/news/vigil-an-agent-harness-for-long-running-pentests-and-code-audits.txt", "jsonld": "https://wpnews.pro/news/vigil-an-agent-harness-for-long-running-pentests-and-code-audits.jsonld"}}