cd /news/ai-safety/depthfirst-continuous-cyber-defense · home › topics › ai-safety › article
[ARTICLE · art-142834] src=depthfirst.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

DepthFirst: Continuous Cyber Defense

DepthFirst launched an AI-native continuous cyber defense product that scans code, dependencies, infrastructure and live environments, and published a benchmark comparing 34 model and thinking-level configurations on detection cost, recall and validation precision. The top configuration, Mythos 5, achieved 69% overall recall at $99.19 in detection cost, while dfs-large1 reached 62.2% recall at $6.77; GPT 6 Sol at xhigh scored 59.3% recall at $13.30 and the lowest-cost configuration, GPT 6 Sol at med, hit 37.3% recall for $0.15. Validation precision ranged from 15.7% to 48.5%, yielding F1 scores between 27.6% and 51.5%.

read8 min views1 publishedSep 30, 2026
DepthFirst: Continuous Cyber Defense
Image: source

AI-native security that thinks like an elite security research team and works across your code, dependencies, infrastructure, and live environment.

%

< min

%

Model and thinking level
Thinking
--- ---
dfs-large1
GPT 5.6 Sol xhigh
GPT 5.6 Sol high
GPT 5.6 Sol med
GPT 5.6 Luna xhigh
GPT 5.6 Luna high
GPT 5.6 Luna med
Opus 5 high
Opus 5 med
Opus 5.5 high
Opus 5.5 med
Grok 4.5 high
Grok 4.7 xhigh
Grok 4.7 high
Kimi K3 max
Kimi K3 high
Qwen 3.8 max
GLM 5.2 xhigh
GLM 5.2 high
DeepSeek v4 Flash max
DeepSeek v4 Flash high
Gemini 3.6 Flash high
Mythos 5*
Gemini 3.8 Flash high
Muse Spark 1.3 xhigh
DeepSeek v4.1 Flash high
GPT 6 Luna xhigh
GPT 6 Luna high
GPT 6 Luna med
GPT 6 Sol xhigh
GPT 6 Sol high
GPT 6 Sol med
MiMo-v2.6 Flash
MiMo-v2.6 Pro
Detection cost and overall recall
$6.77 62.2%
$43.37 65.7%
$26.47 59.6%
$9.51 48.9%
$2.53 52.4%
$1.57 39.6%
$0.69 28%
$22.66 45.4%
$21.89 47.8%
$8.42 54.9%
$7.66 48.1%
$7.70 57.2%
$19.94 59.0%
$16.02 56.6%
$9.10 48%
$6.22 46.7%
$7.91 40.8%
$9.72 40.7%
$4.09 38.2%
$1.23 31.8%
$1.03 30.1%
$7.44 20.9%
$99.19 69%
$8.03 44.2%
$10.24 47.4%
$1.69 57.3%
$0.43 45.7%
$0.28 38.1%
$0.15 37.3%
$13.30 59.3%
$6.45 57.7%
$3.43 49.6%
$0.86 53.5%
$1.38 55%
Detection recall by domain
66.5% 42.4% 62.3%
67.1% 67.3% 60.7%
62.8% 62.3% 49.2%
54.2% 38.5% 44.3%
56.8% 45.1% 47.5%
48.4% 23.1% 31.1%
31.6% 15.4% 29.5%
54.8% 34.1% 31.1%
45.8% 57.7% 44.3%
60.9% — 48.2%
57.5% — 35.9%
59.1% 51.6% 0%
68.4% 42.1% 53.9%
68.4% 15.8% 55.1%
31.8% 31.8% 30.4%
48.4% 51.8% 37.9%
40.8% 40.8% 40.8%
34.8% 22.4% 44.3%
42.6% 11.1% 44.3%
27.6% 29.5% 45.9%
30.1% 30.1% 30.1%
19.4% 11.5% 32.8%
70.8% 61.4% 71%
50.6% 17.3% 50.8%
55.5% 32% 39.3%
64.1% 54.5% 48.9%
50% 40.4% 39.3%
43.9% 25% 34.4%
40.6% 26.9% 37.7%
61.9% 51.9% 59%
58.7% 56.9% 55.7%
48.4% 51.9% 50.8%
54.9% 47.9% 55.3%
58.6% 38.9% 59.8%
Detection recall by repository scope
63% 52.6%
66.3% 57.9%
63.3% 10.5%
50.6% 26.3%
53.2% 42.1%
40.2% 31.6%
29.3% 10.5%
47% 26.3%
48.2% 42.1%
57.2% 31.3%
49.7% 31.6%
59.9% 25%
61.4% 20%
57.7% 40%
49.8% 34.8%
48.5% 31.6%
40.8% 40.8%
43.4% 20.9%
42.2% 36.8%
32.9% 26.3%
31.9% 8.3%
20.9% 21.1%
70.4% 50%
44.8% 36.3%
48.6% 31.6%
58.5% 42.1%
46.4% 36.8%
39.4% 21.1%
38.6% 21.1%
59.4% 57.9%
58.9% 42.1%
50.2% 42.1%
54.3% 28.6%
— —
Validation precision
18.3%
18%
22.3%
24.6%
23.2%
24.8%
27.3%
39.4%
40.5%
48.5%
44.2%
29.5%
23.9%
30.5%
30.1%
33.3%
29.6%
28.6%
30.8%
24.8%
26.8%
43.5%
24.5%
41.4%
34.1%
36.4%
21.8%
24.5%
33.3%
15.7%
18.9%
22.6%
24.4%
18.1%
F1 from detect recall and validate precision
28.3%
28.3%
32.5%
32.7%
32.2%
30.5%
27.6%
42.2%
43.8%
51.5%
46.1%
38.9%
34.0%
39.7%
37.0%
38.9%
34.3%
33.6%
34.1%
27.9%
28.4%
28.2%
36.2%
42.8%
39.7%
44.5%
29.5%
29.8%
35.2%
24.8%
28.5%
31.1%
33.5%
27.2%
Differential analysis
$1.34 75.6%
$10.67 75.3%
$5.46 76.3%
$2.32 74.4%
$0.62 71%
$0.30 71.3%
$0.13 65.1%
$7.41 70.3%
$5.48 71.3%
$1.78 79.5%
$1.34 75.3%
$3.38 65.1%
$6.56 74.8%
$6.46 70.3%
$6.34 60.3%
$4.18 60.1%
$2.99 72.7%
$2.04 61.4%
$1.56 58.6%
$0.79 66.8%
$0.55 67.4%
$1.51 61%
— —
$1.57 63.8%
$3.12 72.8%
$0.71 70.1%
$0.12 72.1%
$0.09 72.4%
$0.07 69.5%
$3.90 75.2%
$2.06 76.3%
$1.00 71.6%
$0.08 72.7%
$0.12 76.5%
Differential analysis by NEW, KEPT, and CLOSED
45.8% 95% 86%
46.8% 95% 84.1%
45.6% 94.4% 88.9%
39.2% 93.3% 90.7%
38.9% 89.3% 84.8%
32.5% 92.4% 89.1%
22.2% 79.2% 94%
38.2% 79.8% 92.8%
37.4% 83.8% 92.8%
66.7% 86.8% 85.2%
47.7% 91.5% 86.8%
39.9% 74.5% 80.9%
51.2% 90.5% 82.6%
43.9% 86.3% 80.8%
35.5% 72% 73.4%
31.4% 75.3% 73.5%
33.6% 95% 89.6%
25.5% 85.6% 73%
25.3% 76.4% 74.1%
28.4% 83% 89%
27.1% 88.1% 87%
18.3% 68.9% 96.5%
— — —
32% 66.3% 93.1%
47.1% 85.1% 86.2%
48.5% 67.5% 94.2%
44.4% 89% 82.9%
45.8% 91.2% 80.1%
36.1% 87.4% 85%
58.3% 92.9% 74.5%
61.8% 92% 75%
52.8% 86.3% 75.9%
38.7% 89.9% 89.5%
55.6% 92.8% 81.3%

Dependency Firewall blocks malicious packages. Security Reviewer validates every human and AI-generated code change before vulnerabilities, sensitive data, or malware enter your codebase.

depthfirst reasons about business logic and cross-service data flows the way a security engineer does, so the only findings you see are the ones that are genuinely exploitable.

depthfirst attacks your running applications continuously, proving at runtime which vulnerabilities are actually exploitable and re-testing every fix after merge.

depthfirst traces every dependency to the code that actually calls it, so you fix the handful of vulnerabilities with a real path into your application instead of the whole SBOM.

Set policy once. Every scan, fix, and agent action inherits it automatically.

Every finding, decision, and fix lives in one place: searchable, exportable, audit-ready.

See who did what, when, and why. Every human and agent action is logged, immutable, and traceable.

Give teams exactly the access they need. No more, no less. Scoped by repo, environment, or org.

Independently audited. Your data handled the way your security team demands.

Bring your own key. Your data stays encrypted under your control, not ours.

── more in #ai-safety 4 stories · sorted by recency
── more on @depthfirst 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/depthfirst-continuou…] indexed:0 read:8min 2026-09-30 · —