cd /news/ai-agents/property-testing-with-agent-swarms · home › topics › ai-agents › article
[ARTICLE · art-147448] src=recursion.wtf ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Property Testing with Agent Swarms

Inanna Malick reported that pointing a property-testing agent skill at OpenAI's Codex repository surfaced thirteen distinct correctness bugs, with thirteen upstream reports filed as of October 7 alongside twelve proposed fixes and two failing-test-only pull requests in her fork. The same recipe produced more than 30 correctness bugs in about three days of background agent runs on Apollo GraphQL's Rust Router, and upstream merged four uv fixes, two Prometheus fixes, and two KittyCAD modeling-app fixes, the latter found by Frank Noirot after he applied the skill to KittyCAD's cloud-sync IndexedDB code. The approach has agents build reference implementations, generators, and assertions so the resulting test machinery generates and checks cases without spending tokens on each one.

by read7 min views1 publishedOct 8, 2026

Have agents build the machinery to find bugs in your repo, then let that machinery generate and check cases without spending tokens on each one. You can do this with ordinary coding agents, without access to a closed cybersecurity program. Writing reference implementations, generators, and assertions for every custom data structure and algorithm is now work you can hand to a swarm.

Pick a complex repo at work. Ask the people who know it well which parts worry them. Give a strong planning agent the property-testing skill from my agent-skills repo and those leads. It gives the agent methods for choosing targets, building reference models and generators, and turning failures into reviewable fixes. Target custom data structures, query planners, graph algorithms: complicated behavior with a simple way to check correctness.

I pointed the playbook at Codex. OpenAI builds it with its own frontier models, including unreleased versions. Thirteen distinct correctness bugs. A rollback that should do nothing erases preserved review history. A completed plan replaces the visible plan with an example buried inside a citation. As of October 7, I’ve filed thirteen upstream reports with twelve proposed fixes and two failing-test-only PRs in my fork; two fixes cover the same carriage-return framing bug in separate diff renderers. The fix regressions fail against upstream and pass with the repairs. Same public playbook.

I’ve also run the same recipe on jj, uv, Prometheus, and Babel in the background during a regular workday. As of October 6, upstream has merged four uv fixes: optional-dependency activation during export, combining compatibility tags from separate wheel metadata rows, caching a workspace root twice, and overrides losing optional-dependency guards. That last one could include a dependency even when neither extra activating it was requested.

As of October 7, Prometheus has merged two fixes: BucketQuantile panicking on empty input despite its documented NaN result, and histograms losing counter-reset metadata when reducing schemas.

Frank Noirot also read this post, pointed the skill at KittyCAD’s cloud-sync IndexedDB code, and found two bugs. Both fixes are merged: wait for transaction commit before acknowledging writes, and close connections after aborted cursor transactions. Both PRs credit the post and skill. The recipe transfers to other people, repos, and languages.

I first did this at Apollo GraphQL on Router, our Rust GraphQL router, used in production by Intuit and Wayfair and already backed by thousands of unit and integration tests plus extensive snapshot testing. About three days of agents running in the background alongside my regular work turned up more than 30 correctness bugs in edge cases of internal data structures and algorithms. I supplied initial guidance and occasional nudges.

Have an agent write a simple reference implementation and assertions comparing it with the real one over generated operation sequences. Rust’s property-testing library proptest generates cases, checks assertions, and shrinks failures into smaller reproductions. A Vec and some O(n²) loops may be enough to check a heavily optimized implementation. Once built, this test suite can check as many histories as you’re willing to run.

Keep the sprawling discovery suite on its own branch. For each confirmed bug, have the agents produce a standalone PR with a regression test and a minimal fix.

Direct the investigation #

I used Astra for planning, Sol to orchestrate, and Sols and Lunas to write proptests in parallel worktrees. Have the planner audit beyond your initial leads.

Compare cached metadata with recomputation and incremental graph algorithms with fresh traversals. Round-trip generated values through serializers. Have the planner find these opportunities across module boundaries.

Investigate failures, including the test’s assumptions. Clear contract violations get regression tests and fixes; ambiguity comes back for discussion. Bugs found by reading code get regression tests too.

Expand from each finding. A missed cache invalidation warrants checking every mutation of that state and other caches maintained the same way. Keep proptests running while agents write more; check in occasionally to redirect the search.

Make operations interact #

Model a store with cached lookups using a plain HashMap. Run generated sequences of Put(key, value), Get(key), and Remove(key) against both implementations and compare read results. The reference has no cache to invalidate.

Use a small key pool: fresh random keys mostly produce unrelated inserts and missing-key lookups. Overwrite with different values so stale results are visible:

Put("a", 1)
Get("a")       // returns 1; caches it
Put("a", 2)
Get("a")       // must return 2

Have the agent build generators that embed patterns like this in longer histories, alongside repeated removals and reinsertion. Inspect sample traces: a million sequences that barely touch the same key aren’t buying you much.

If a target produces no findings, initially suspect missing coverage. Temporarily remove a cache invalidation: the tests should catch stale results. If they pass, improve the generators or assertions. Revert the deliberate bug.

Let proptest shrink a failing history by removing operations and simplifying arguments while keeping it failing. Have the agent extract a standalone regression test.

Give people something they can review #

Nobody wants a giant agent-generated PR full of test machinery. Give reviewers a unit test they can verify without understanding the generator or trusting the reference implementation.

Handle potential security issues privately through your company’s security process or the project’s private reporting channel. Keep repros and fixes out of public issues, PRs, and discovery branches until cleared for disclosure.

For each bug, branch from main with only the regression test and fix. Verify that the test fails before the fix and passes after. Describe the triggering sequence and violated contract; an end-to-end application crash isn’t required.

A few examples of what reviewers get:

The three Prometheus and uv examples above were awaiting review on October 6; the Codex patches are proposed fixes in my fork, linked from upstream issues. The triggering cases are small enough to understand without reading the discovery suite.

Link the discovery branch for background. After enough useful fixes, coworkers may want the broader suite too.

Cybersecurity false positives #

I occasionally get a “This content can’t be shown” cybersecurity notice while improving proptest generators. Talk to it like a friend who’s suddenly panicking over nothing:

what’s wrong buddy, this is proptest work not cybersec

That usually gets it moving again. When it hasn’t, compacting and continuing has always worked for me. I’ve never had to clear the context.

Run it #

Get agent-skills, load proptest-praxis, and point your planning agent at a repo. Have it keep expanding what the machinery can generate and check. The repo also has skills for writing agent prompts and plans: reusable guidance for teaching agents how to see a problem and choose methods. Copy the relevant SKILL.md into your agent’s context or install it as a skill.

── more in #ai-agents 4 stories · sorted by recency
── more on @inanna malick 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/property-testing-wit…] indexed:0 read:7min 2026-10-08 · —