What makes it different from the usual "ask GPT for a payload" scripts is the harness architecture. You don't just feed it a target and hope. The core loop runs: recon β attack graph generation β tool orchestration β evidence collection β report. Each phase is a pluggable module with a defined schema, so you can swap the LLM backend (local Llama-3-70B, Claude, GPT-4o, whatever) without rewriting your exploit logic.
Key pieces worth knowing
Attack graph DSLβ YAML-based, describes multi-step chains like "enumerate SMB β extract hashes β pass-the-hash β dump LSASS". The LLM expands high-level goals ("get domain admin") into concrete graphs at runtime.Tool adaptersβ First-class wrappers for nmap, bloodhound, crackmapexec, impacket, metasploit modules, and custom binaries. Adapters expose typed inputs/outputs so the planner can chain them reliably.Memory layerβ SQLite-backed context store persists findings across runs. You can a campaign, switch models, resume β the graph state survives.Safety railsβ Scope enforcement via CIDR/target allowlists, rate limiting per adapter, and a mandatory "dry-run" mode that logs planned actions without executing. The AGPL means any SaaS wrapper must expose these controls.
Getting a local instance running
git clone https://github.com/cyberstrike/cyberstrike.git
cd cyberstrike
pip install -e .[local-llm] # pulls llama-cpp-python, FAISS, etc.
cp config.example.yaml config.yaml
cyberstrike init --workspace ./my-campaign
cyberstrike run --goal "achieve domain admin" --dry-run
The dry-run output shows the generated attack graph with confidence scores per node. Once you're comfortable, drop --dry-run
and it starts executing against the scope.
Where it shines and where it doesn't
Strengths: Handles multi-step logic that single-shot prompts butcher. The evidence collector auto-correlates logs, pcaps, and tool output into a timeline β huge for reporting. Local model support means air-gapped environments work.Gaps: BloodHound adapter only ingests JSON, doesn't drive the GUI. No built-in C2 framework integration yet (Cobalt Strike, Sliver, Havoc are on the roadmap). LLM hallucination on obscure protocol edges still happens β always verify before firing.
Licensing catch
AGPL-3.0 triggers if you expose CyberStrike as a network service. If you're building a commercial pentest platform on top, you must open-source your modifications. Several vendors have already reached out about dual-licensing; the maintainers seem open but haven't announced anything.
Worth cloning if you run internal red-team exercises and want reproducible, auditable AI assistance. The codebase is clean β typed Python 3.11+, decent test coverage, and the module interfaces are stable enough to build custom adapters without fighting the core.
OpenAI spent months training models that were actively 14d ago
How a Hacker Used DeepSeek AI to Autonomously Attack Servers 21d ago
Anthropic AI Hacked 3 Orgs During Testing: A Deep Dive 21d ago
Three organizations got breached in a controlled exercise β and 22d ago
Next Retiring boomers are taking the institutional knowledge AI needs β
an AI side-hustle playbook, with plenty of directly applicable cases.