cd /news/artificial-intelligence/ai-confidence-is-high-evidence-lags-… · home › topics › artificial-intelligence › article
[ARTICLE · art-142743] src=sdtimes.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI confidence is high, evidence lags behind, SmartBear report finds

SmartBear's State of Software Quality and Testing 2026 report found that 81% of organizational leaders believe AI can reliably catch its own errors despite having no evidence, while only 64% of software practitioners trust it. The survey also found 46% of teams shipped AI-generated code that later failed, yet 69% of those teams still reported a lot or complete confidence the code acted as created, and only 3% of respondents rely solely on AI self-validation while 84% use at least one form of human review. The report recommends validating specifications before AI generates code, independent review of more than 60% of AI-generated code before commit (49% of respondents) and before release (60%).

by read3 min views1 publishedSep 30, 2026
AI confidence is high, evidence lags behind, SmartBear report finds
Image: Sdtimes (auto-discovered)

While 81% of organizational leaders believe AI can reliably catch its own errors, despite having any evidence, only 64% of software practitioners trust it, exposing a gap in confidence about AI software quality, according to SmartBear’s State of Software Quality and Testing 2026 report. In fact, according to the survey, 46% of teams have shipped AI-generated code that later failed, but of those, 69% still say that have a lot or complete confidence that the code is acting as it was created. The divide between what leaders believe and what practitioners see is wide, and it is up to the developers and engineers working with the software to make their results match the expectations of leadership, “whether or not they have the proper tools and processes to do so,” the report noted.

Validation has never been more important

While the use of AI to test software is needed as AI generates more code at great speed, organizations need the right checks in place to verify its work throughout the development life cycle. Part of that validation includes having people checking AI’s work to ensure the software works as intended and is not riddled with errors. And again, developers are more cautious about AI-generated tests than leaders, with 56% reporting they trust AI tests more than human-written tests. Meanwhile, 72% of leaders have that trust. But organizations aren’t quite ready to abandon testers, as only 3% of respondents say they rely solely on AI self-validation, while 84% use at least one form of human review for validation of AI-created tests, the report found. Yet interestingly, 92% of respondents said AI is the primary tester of its own code, and 48% say it’s only appropriate when paired with human oversight.

One of the key arguments against having AI test its own code is that reviewers cannot validate that code if they don’t know what specification the AI is writing against, SmartBear wrote in the report. “Having the same system responsible for generating the code and validating it results in a testing black box – neither the logic behind the tests nor the assumptions baked into them are independently legible to the humans nominally overseeing the process,” according to the report. “Teams have no separate signal to confirm that what is being tested reflects what actually matters, and no clear line of sight into whether a passing test suite represents genuine coverage or simply AI confirming its own work.”

Stages of independent checks

Historically, teams would check code the closer it gets to shipping, but when it became important to ship more frequently, the “shift-left” movement began to start testing earlier in the development life cycle. Now, with AI, organizations are employing a three-stage approach to testing:

Before generation: 46% of responders say they check that the specificaiton reflects the intent of the code before AI generates it.

Before commit:  49% report an independent review ofmore than 60% of AI-generated code before commiting.

Before release: 60% say the independently test more than 60% of the AI-written code before releasing it.

“Validating the specification before AI ever starts coding from it is the earliest and cheapest point to catch a problem.” the report said. “Bugs found later in the SDLC are the most costly and time-consuming to fix. So while AI is helping some testing teams achieve more coverage, it’s not necessarily solving the problem of cost. Thorough testing is more than just catching the issues; it’s doing so before they become a drain on the work.”

To learn more, read the full report here.

── more in #artificial-intelligence 4 stories · sorted by recency
github.com · · #artificial-intelligence
AI Blinder
── more on @smartbear 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-confidence-is-hig…] indexed:0 read:3min 2026-09-30 · —