cd /news/artificial-intelligence/validating-ai-models-and-agents-with… · home › topics › artificial-intelligence › article
[ARTICLE · art-142388] src=infoworld.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Validating AI models and agents with property-based testing

Property-based testing (PBT) offers a way to validate stochastic AI models and agents by asserting invariants across a distribution of inputs rather than checking single expected outputs, according to an InfoWorld article citing Alan LeFort, CEO and cofounder at StrongestLayer. The technique, which dates to 1997 security research and a 2000 paper on formulating and testing properties, addresses the oracle problem — where no practical way exists to determine the correct classification for ambiguous inputs such as a "Marketplace $97.63" line on a hotel folio. A PBT generator can create hundreds of receipts tested several times each, plus thousands of meaning-preserving variants, to find where semantically identical inputs yield different results and expose defects.

read8 min views5 publishedSep 30, 2026

Testing deterministic systems is relatively straightforward. Create an assertion that the system should pass, and automate validating it against a series of input-to-output data patterns. When new test patterns are needed, use observability sources to extract them from actual usage data or to create synthetic test data. But testing AI models and agents with nondeterministic outputs doesn’t align with this form of unit testing. For example, say you are using a language model to categorize a string of text, such as text extracted from an employee’s expense receipts. If the number of categories is small and there is little variance between employee receipts, then traditional testing methods can work when you have sufficient test cases. But this approach doesn’t work if there are hundreds of categories and an untold number of receipt variants.

In my 2024 article on testing large language models, I reviewed several testing approaches, including using adversarial networks for evaluating model output. In my 2025 article on testing AI agents, I covered testing strategies and approaches for validating response accuracy. In the present article, I dive deeper into property-based testing (PBT), an approach that uses property specifications to evaluate tests for correctness and completeness. PBT evaluates a system using properties rather than expected outputs. The term dates from 1997 security research, and the technique stems from a 2000 paper on formulating and testing properties.

“When the system under test is stochastic, point assertions stop being useful because the same input can produce different outputs, so traditional pass-fail testing breaks down,” says Alan LeFort, CEO and cofounder at StrongestLayer. “Property-based testing flips the problem, and instead of asserting one expected result, you assert an invariant that should hold across a distribution of inputs and conditions.”

AI models and agents are stochastic, and even identical input prompts can generate different results. PBT’s strength is that it doesn’t need labeled test data for testing and can generate the input patterns. The test’s objective is to find where semantically identical inputs yield different results, and its strength is in helping identify which input patterns make the model unstable.

PBT works by creating an assertion, such as “each line item in a receipt is assigned to one category,” and then identifying the list of categories, such as meals, grocery, airfare, and hotel expenses. That’s the property being tested; the categories are the allowed outputs, and the generator is the engine that creates the inputs.

The hard problem PBT solves for is that you often can’t say what the correct category is for ambiguous inputs. A line reading “Marketplace $97.63” on a hotel folio could be classified as hotel, meals, or grocery. Testing researchers call this the oracle problem, where there’s no practical way to determine the correct classification.

A PBT generator can be applied by creating hundreds of receipts, each tested several times to measure when the model is consistent. This is one way to validate a model’s stability. Then, create thousands of variants of each receipt that don’t change the underlying meaning; for example, reordering the lines, reformatting the amounts, or changing the guest name. If the same line gets classified two ways for inputs that mean the same thing, at least one is wrong, and you’ve found a defect even though you never learned which classification was correct.

PBT takes inputs that created defects and runs a systematic search called shrinking. It simplifies the input by dropping line items, reducing amounts, and shortening descriptions, then reruns the simplified receipt and its variant through the assertion. It keeps only simplifications that still fail by returning different results, and repeats the shrinking and asserting process until no further simplification is possible while still failing.

What comes back is the smallest input that still breaks. A hotel folio with multiple line items might shrink to two lines: “Room Charge $289.00” and “Marketplace $97.63.” Reordering those two lines flips Marketplace from hotel to grocery, highlighting the defect. The minimal case tells you where the model is unstable; here, a line’s category depends on what precedes it.

“Property-based testing shifts AI validation from checking whether one answer is correct to proving that critical behaviors never fail, regardless of prompts, inputs, or enterprise context,“ says Vijay Rayapati, cofounder and CEO at Atomicwork. “For enterprise AI agents, those properties include never granting unauthorized access, bypassing approvals, exceeding permissions, or taking unsafe actions. Running thousands of automatically generated scenarios exposes edge cases that traditional testing misses.”

Sanmi Koyejo, cofounder and head of AI at Virtue AI, shared two examples of how PBT can be used to validate an AI language model: paraphrasing a prompt shouldn’t flip a safety verdict, and an AI agent shouldn’t cross a permission boundary no matter how the request is worded. “Property-based testing is strong where you can state an invariant that must hold across a huge input space. The generator finds edge cases your example tests never would, and shrinking hands you a minimal reproduction. In practice, lean on metamorphic properties such as consistency and monotonicity because there’s no exact-match oracle for generative output and seed generators with realistic traffic plus adversarial mutations,” says Koyejo.

Two other practical examples where PBT has been applied include the inverse application of AI writing the tests rather than being tested by them, and validating that an AI agent sticks to its area of expertise.

Mohammed Aboul-Magd, vice president of product for the cybersecurity group at SandboxAQ, says the most dangerous agent failures are quiet boundary violations. “When a support agent answers questions it was never meant to handle, or a retrieval agent accesses data outside its scope, an organization is at risk. PBT catches those before production,” Aboul-Magd says.

Traditional testing often targets what’s called the known knowns, their testing patterns, and boundary cases. PBT helps extend testing into unknown areas where you have no reliable or valid data to test, or little understanding of which conditions yield unexpected or risky results.

“The most expensive AI bugs rarely show up in the examples you expected people to try,” says Zain Lakhani, chief AI officer at Pendo. “Property-based testing pushes models and agents through thousands of unexpected scenarios, exposing behaviors that scripted tests often overlook. That extra coverage helps teams catch reliability issues before users do.”

Another reason experts explore PBT is to generate test cases for AI code generators and for applications developed with vibe coding or spec-driven development practices. Relying on unit tests generated by code generators may lead to “happy path, known knowns” testing patterns.

“Property-based testing gives teams a way to define what ‘correct’ means and then aggressively pressure-test that across thousands of inputs,” says Ali-Reza Adl-Tabatabai, CEO and cofounder at Gitar. “That matters when AI can generate and modify code faster than humans can review it. But testing alone won’t keep up.”

Harshil Shah, director of engineering and head of the AI center of excellence at R Systems, RSI,  says that the teams getting real value from property-based testing of agents make four implementation moves:

PBT can also be used to validate a language model’s invariants, the structural or semantic properties that must hold for any valid input. These include testing output length, schema validation, and consistency under rephrasing.

“Property-based testing adds value to the AI validation toolkit by shifting the focus from checking individual outputs to verifying that important behavioral properties hold across a wide range of scenarios,” says Sridhar V. Dasaratha, EY distinguished technologist and EY global client technology AI research leader at EY. “For AI models and agents, it can help uncover edge cases and validate critical guardrails such as safety, compliance, workflow integrity, and responsible tool use. A key best practice is to define clear, business-relevant properties that can be independently verified and monitored.”

PBT is a promising testing approach, but it does have some implementation considerations and limitations. There are costs in using LLM-based generators. Engineers must also understand the trade-offs between testing enough variants to trust the results and the time and cost to execute the tests. Some areas may be blind spots if the team doesn’t know they need to test for them. Koyejo of Virtue AI adds, “PBT only checks properties you thought to write, so it misses failure modes you didn’t anticipate. Pair it with adversarial red teaming to surface the unknown unknowns it can’t reach.”

Scaling PBT may be a challenge, especially as AI agents become more sophisticated and connect to AI orchestration platforms to support complex workflows. Joseph Hurley, principal sales engineer at Digital.ai, says, “Proving a property that holds for a handful of functions is one thing, but proving rigorous quality standards that hold across an entire AI-generated system is another.”

PBT is only one validation step, and teams should determine their AI agents’ release-ready criteria, with additional steps to validate security, measure business value, and consider AI cost debt.

“Property-based testing is a powerful way to formalize expectations for AI systems and creates a strong foundation for specifying how systems should behave across a wide range of inputs,” says Bo Li, CEO and cofounder at Virtue AI. “But true validation goes beyond abstract properties and requires observing agents in realistic environments with live tools, dynamic permissions, and adversarial pressures that expose emergent risks.”

As more organizations develop AI agents, using PBT and other testing methods will be a key practice for deploying reliable, trustworthy agents.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @infoworld 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/validating-ai-models…] indexed:0 read:8min 2026-09-30 · —