Property Testing with Agent Swarms Inanna Malick reported that pointing a property-testing agent skill at OpenAI's Codex repository surfaced thirteen distinct correctness bugs, with thirteen upstream reports filed as of October 7 alongside twelve proposed fixes and two failing-test-only pull requests in her fork. The same recipe produced more than 30 correctness bugs in about three days of background agent runs on Apollo GraphQL's Rust Router, and upstream merged four uv fixes, two Prometheus fixes, and two KittyCAD modeling-app fixes, the latter found by Frank Noirot after he applied the skill to KittyCAD's cloud-sync IndexedDB code. The approach has agents build reference implementations, generators, and assertions so the resulting test machinery generates and checks cases without spending tokens on each one. Have agents build the machinery to find bugs in your repo, then let that machinery generate and check cases without spending tokens on each one. You can do this with ordinary coding agents, without access to a closed cybersecurity program. Writing reference implementations, generators, and assertions for every custom data structure and algorithm is now work you can hand to a swarm. Pick a complex repo at work. Ask the people who know it well which parts worry them. Give a strong planning agent the property-testing skill https://github.com/inanna-malick/agent-skills/blob/main/skills/proptest-praxis/SKILL.md from my agent-skills repo https://github.com/inanna-malick/agent-skills and those leads. It gives the agent methods for choosing targets, building reference models and generators, and turning failures into reviewable fixes. Target custom data structures, query planners, graph algorithms: complicated behavior with a simple way to check correctness. I pointed the playbook at Codex https://github.com/openai/codex . OpenAI builds it with its own frontier models, including unreleased versions. Thirteen distinct correctness bugs. A rollback that should do nothing erases preserved review history https://github.com/openai/codex/issues/51748 . A completed plan replaces the visible plan with an example buried inside a citation https://github.com/openai/codex/issues/51773 . As of October 7, I’ve filed thirteen upstream reports with twelve proposed fixes and two failing-test-only PRs in my fork https://github.com/inanna-malick/codex/pulls?q=is%3Apr ; two fixes cover the same carriage-return framing bug in separate diff renderers. The fix regressions fail against upstream and pass with the repairs. Same public playbook. I’ve also run the same recipe on jj, uv, Prometheus, and Babel in the background during a regular workday. As of October 6, upstream has merged four uv fixes: optional-dependency activation during export https://github.com/astral-sh/uv/pull/22234 , combining compatibility tags from separate wheel metadata rows https://github.com/astral-sh/uv/pull/22235 , caching a workspace root twice https://github.com/astral-sh/uv/pull/22236 , and overrides losing optional-dependency guards https://github.com/astral-sh/uv/pull/22237 . That last one could include a dependency even when neither extra activating it was requested. As of October 7, Prometheus has merged two fixes: BucketQuantile panicking on empty input despite its documented NaN result https://github.com/prometheus/prometheus/pull/19927 , and histograms losing counter-reset metadata when reducing schemas https://github.com/prometheus/prometheus/pull/19916 . Frank Noirot also read this post, pointed the skill at KittyCAD’s cloud-sync IndexedDB code, and found two bugs. Both fixes are merged: wait for transaction commit before acknowledging writes https://github.com/KittyCAD/modeling-app/pull/14396 , and close connections after aborted cursor transactions https://github.com/KittyCAD/modeling-app/pull/14397 . Both PRs credit the post and skill. The recipe transfers to other people, repos, and languages. I first did this at Apollo GraphQL on Router https://github.com/apollographql/router , our Rust GraphQL router, used in production by Intuit https://www.apollographql.com/blog/how-intuit-handled-their-busiest-time-of-year-with-apollo-router and Wayfair https://www.apollographql.com/events/how-wayfair-slashed-costs-simplified-infra-and-cut-latency-in-half-with-apollo-router and already backed by thousands of unit and integration tests plus extensive snapshot testing. About three days of agents running in the background alongside my regular work turned up more than 30 correctness bugs https://github.com/apollographql/router/pulls?q=is%3Apr%20author%3Ainanna-apollo%20created%3A2026-09-01..2026-10-02%20-head%3Ainanna%2Fgraph-proptest-coverage in edge cases of internal data structures and algorithms. I supplied initial guidance and occasional nudges. Have an agent write a simple reference implementation and assertions comparing it with the real one over generated operation sequences. Rust’s property-testing library proptest https://proptest-rs.github.io/proptest/intro.html generates cases, checks assertions, and shrinks failures into smaller reproductions. A Vec and some O n² loops may be enough to check a heavily optimized implementation. Once built, this test suite can check as many histories as you’re willing to run. Keep the sprawling discovery suite on its own branch. For each confirmed bug, have the agents produce a standalone PR with a regression test and a minimal fix. Direct the investigation I used Astra for planning, Sol to orchestrate, and Sols and Lunas to write proptests in parallel worktrees. Have the planner audit beyond your initial leads. Compare cached metadata with recomputation and incremental graph algorithms with fresh traversals. Round-trip generated values through serializers. Have the planner find these opportunities across module boundaries. Investigate failures, including the test’s assumptions. Clear contract violations get regression tests and fixes; ambiguity comes back for discussion. Bugs found by reading code get regression tests too. Expand from each finding. A missed cache invalidation warrants checking every mutation of that state and other caches maintained the same way. Keep proptests running while agents write more; check in occasionally to redirect the search. Make operations interact Model a store with cached lookups using a plain HashMap . Run generated sequences of Put key, value , Get key , and Remove key against both implementations and compare read results. The reference has no cache to invalidate. Use a small key pool: fresh random keys mostly produce unrelated inserts and missing-key lookups. Overwrite with different values so stale results are visible: Put "a", 1 Get "a" // returns 1; caches it Put "a", 2 Get "a" // must return 2 Have the agent build generators that embed patterns like this in longer histories, alongside repeated removals and reinsertion. Inspect sample traces: a million sequences that barely touch the same key aren’t buying you much. If a target produces no findings, initially suspect missing coverage. Temporarily remove a cache invalidation: the tests should catch stale results. If they pass, improve the generators or assertions. Revert the deliberate bug. Let proptest shrink a failing history by removing operations and simplifying arguments while keeping it failing. Have the agent extract a standalone regression test. Give people something they can review Nobody wants a giant agent-generated PR full of test machinery. Give reviewers a unit test they can verify without understanding the generator or trusting the reference implementation. Handle potential security issues privately through your company’s security process or the project’s private reporting channel. Keep repros and fixes out of public issues, PRs, and discovery branches until cleared for disclosure. For each bug, branch from main with only the regression test and fix. Verify that the test fails before the fix and passes after. Describe the triggering sequence and violated contract; an end-to-end application crash isn’t required. A few examples of what reviewers get: - Codex root snapshots https://github.com/inanna-malick/codex/pull/10 : replayed copies of one assistant message consume the shared message limit, crowding out three of eight distinct user messages in a persisted-and-resumed conversation’s root snapshot. - Codex inline-tag parsing https://github.com/openai/codex/issues/51723 : configure opening delimiters