cd /news/artificial-intelligence/ai-pentesting-evaluation-checklist-w… · home topics artificial-intelligence article
[ARTICLE · art-103775] src=aikido.dev ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI pentesting evaluation checklist: What to look for in an AI pentesting vendor

Aikido's State of AI in Pentesting report, surveying 200 CISOs and 200 engineering leaders, found that 76% ship significant changes weekly or faster but only 21% validate security on every release. The company's evaluation checklist emphasizes that AI pentesting vendors should be judged on validation quality, depth of testing, and code access, noting that engagements with code access surfaced a median of 7x more high and critical vulnerabilities than tests without it. In one case, a financial services platform's 120-hour manual pentest returned zero findings, while AI pentesting agents with codebase access took just over two hours and found 13 valid issues, including three high-severity findings.

read7 min views3 publishedAug 20, 2026
AI pentesting evaluation checklist: What to look for in an AI pentesting vendor
Image: Aikido (auto-discovered)

AI pentesting is still a relatively new category, and the term covers very different products dressed in the same language. Some tools behave more like automated scanners (DAST with LLM lipstick), while others explore application workflows, attempt multi-step attacks, and validate findings against a live system. That makes it important to look beyond the terminology and evaluate what the product actually does in practice.

In our State of AI in Pentesting report, we surveyed 200 CISOs and 200 engineering leaders. 76% said they ship significant changes weekly or faster, but only 21% validate security on every release. That is what AI penetration testing is meant for.

This checklist covers the areas we think matter most when evaluating an AI pentesting vendor, including validation and signal quality, depth of testing, code access, operational reliability, onboarding, technical guardrails, model flexibility, and how well testing fits into your existing development workflow.

A simple way to make the eval more meaningful is to request direct access from each vendor and watch the test run live, rather than comparing final reports alone. Keep the testing scope comparable, understand how much effort was required to get each test running, and account for how much budget or testing capacity each vendor consumed. That gives you a much closer apples-to-apples comparison.

AI pentesting evaluation criteria at a glance #

  • Validation and signal quality: findings come with proof you can act on directly
  • Depth of AI pentesting: agents test permission boundaries and multi-step flows across the application
  • Visibility and reporting: you can see what was tested and how
  • Speed and workflow fit: results arrive in time to be relevant
  • Operational reliability: tests complete without constant supervision
  • Scope control and safety: enforcement happens in the system itself
  • Setup and onboarding: a real test runs without a long ramp-up
  • Model flexibility: not locked to a single model or provider

AI pentesting evaluation checklist #

Most vendors will look strong in a demo. The following criteria focus on how their product performs in practice.

Validation and signal quality

  • Findings include clear proof and can be reproduced
  • Issues are validated against a live system
  • Results can be used directly without additional triage
  • Findings reflect actual exploit paths

Depth of AI pentesting

  • Testing covers workflows, roles, and application state
  • Multi-step issues can be discovered and confirmed
  • Testing adapts to how the application behaves
  • Can it use source code or internal context to improve coverage?

Code access changes the outcome here more than any other factor on this list. Across more than 1,000 AI-driven penetration tests, engagements with code access surfaced a median of 7x more high and critical vulnerabilities than tests without it, at roughly half the agent cost per finding.

Visibility and reporting

  • You can see what was tested and how the system responded
  • There is a clear record of actions taken during testing
  • Reports contain enough detail to reproduce issues

Speed and workflow fit

  • A meaningful test can be run without long setup
  • Results arrive in time to act on them
  • Fixes can be retested and confirmed
  • Testing fits into how your team releases software

How fast a test runs affects how much of the system it actually reaches. One financial services platform ran a 120-hour manual pentest that returned zero findings. The same application, tested with AI pentesting agents that had codebase access, took just over two hours and surfaced 13 valid issues, including three high-severity findings the manual engagement missed entirely. Ask vendors to show you comparable timing alongside the final report.

Operational reliability

  • Tests complete without frequent restarts

  • Failures are understandable when they occur

  • Runs do not require constant supervision

  • The system works consistently in day-to-day use Every system will fail occasionally. What matters is how often it happens and how much effort is required to recover.

Scope control and safety

  • Scope is enforced during execution
  • Scope restrictions are enforced in the system itself, independent of instructions or policy
  • Access is limited to defined domains
  • Production is only included when explicitly configured
  • Out-of-scope requests are blocked automatically
  • Pre-flight checks validate scope and environment configuration before testing begins

Ask how a vendor handles the agents that don't behave predictably. As Aikido's AI pentest lead Philippe Dourassov puts it: "There's going to be five percent of agents that aren't always sensible, and that's why we make sure that we deal with this five percent." A vendor without a clear answer here is depending on the model's judgment instead of a system-level guardrail.

For more information on securing AI pentesting agents by design, read this.

Isolation and data protection

  • Outbound access and data exfiltration are restricted during execution
  • Data stays within approved environments
  • Execution environments are isolated from each other
  • Tests can be monitored, d, and terminated immediately

Setup and onboarding

  • A test can be started without a long onboarding process
  • Configuration is straightforward
  • Authentication and scope setup are stable

Setup is often where problems show up early. If getting started is difficult, it usually carries through into regular use.

Model flexibility and validation depth

  • The product is not locked to a single AI model or provider
  • Different models can be used for different testing tasks
  • Multiple agents can explore and test simultaneously
  • Agents can adapt based on application behavior during testing
  • Testing is not limited to predefined vulnerability signatures
  • Agents can follow multi-step attack paths and chain weaknesses together
  • Findings are confirmed by running them against a live system, so nothing is inferred from model output alone
  • Model behavior is constrained through technical guardrails during execution

If a vendor cannot clearly meet these criteria, the results will eventually require manual validation to be trusted, which defeats the purpose of AI pentesting. Red flags to watch for

  • Tests don't complete reliably
  • Setup takes too long
  • Findings aren't proven
  • Results don't lead to fixes
- No built-in retesting
- Only tests surface-level behavior
  • Scope isn't enforced technically
  • Looks good in a demo, struggles in practice

Quick list to tick off from vendors

  • Runs when you ship
  • Has the option of code access for deeper coverage
  • Finds real, validated issues
  • Catches logic and authorization flaws
  • Easy to get running
  • Stays within defined scope
  • Is fast enough to fit your release schedule
  • Results you can reproduce
  • Fixes get verified
  • Fits how your team ships and fixes code
  • Runs reliably

Proof from the field

Tyro, the Australian financial payments firm, ran a proof of concept comparing AI pentesting against its incumbent human testers on a new health product before launch.

"Our incumbent penetration testers took approximately 15 days. We did the same thing with Aikido's AI penetration testing. It took five and a half hours and found approximately 30 issues, nine high, whereas the human pen testers found five issues, one high,” said Tyro’s former information security officer, Arun Singh.

The plan going forward keeps the human reviewers in place and adds AI to widen coverage: "Our plan is a hybrid model, AI penetration testing and AI code audit with a human review from our pentesting partners. The costs are significantly less than what we were paying, so we can do a lot more with what we have."

Put it to the test #

The fastest way to apply this checklist is to see how the vendors work themselves. Request platform access, run a comparable test with each one, and score them against the sections above while the test is running.

See how these criteria hold up on a live platform, or read the full Buyer's Guide for AI Pentesting for the complete research behind this checklist.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @aikido 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-pentesting-evalua…] indexed:0 read:7min 2026-08-20 ·