cd /news/developer-tools/one-command-to-verify-an-agent-patch… · home topics developer-tools article
[ARTICLE · art-117387] src=dev.to ↗ pub= topic=developer-tools verified=true sentiment=· neutral

One Command to Verify an Agent Patch: Seeded Hypothesis, Fixture Contracts, and a TTL Quarantine

MonkeyCode's outreach article presents a three-rung verification ladder for agent-generated patches in Python projects, combining seeded Hypothesis property tests, fixture contracts, and a TTL-based quarantine for flaky tests. The approach runs under a single pytest command, ensuring deterministic verification without trusting model-written tests blindly.

read5 min views1 publishedSep 1, 2026

A merge gate should be a single command, not a mental checklist. When the patch comes from an agent, the gate matters even more, because you did not watch the reasoning happen. The patch is a hypothesis; your verification should be a repeatable experiment.

This post walks through a three-rung ladder you can add to a Python project:

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

All three rungs run under one command and produce standard pytest output. No generated test is trusted just because a model wrote it; it has to pass the same deterministic harness as any hand-written test. The only dependencies are pytest

and hypothesis

.

Example-based tests answer the question: what did this input do? Property tests answer a different question: what is never allowed to happen? A generated patch often changes behavior at a boundary that no example covers, so a fixed set of examples is the wrong oracle.

Here is a minimal property test for a discount helper:

from hypothesis import given, settings, strategies as st

SEED = 20260901

def discount(price, rate):
    if not 0.0 <= rate <= 1.0:
        raise ValueError('rate out of range')
    return round(max(0.0, price * (1.0 - rate)), 2)

@given(
    st.floats(min_value=0.0, max_value=1_000_000.0),
    st.floats(min_value=0.0, max_value=1.0),
)
@settings(max_examples=300, seed=SEED, deadline=None)
def test_discount_invariants(price, rate):
    result = discount(price, rate)
    assert result >= 0.0
    assert result <= price + 1e-6
    if rate == 0.0:
        assert result == round(price, 2)

The seed matters. If Hypothesis finds a counterexample, you can rerun the exact same search when you fix the code. Without a seed, a flaky property test becomes part of the flake problem you are trying to solve.

A property test can also catch a regression that a single unit test misses. For instance, a patch that computes round(price - price * rate, 2)

may pass a table of hard-coded examples and still break the monotonicity expectation at extreme input scales. The property test turns that invisible break into a normal CI failure.

Fixtures are not just data. They are a contract between the test and the domain. If an agent patch modifies the fixture , a changed field name will cause a confusing failure deep inside the suite. A fixture contract catches that before any domain assertion runs.

import json
from pathlib import Path

CONTRACTS = {
    'ledger_entry': {
        'required': ('id', 'amount', 'currency'),
        'semantic_checks': (
            lambda row: row['amount'] > 0,
            lambda row: row['currency'] in ('USD', 'EUR'),
        ),
    }
}

def test_fixture_contracts():
    for path in Path('tests/fixtures').glob('*.json'):
        row = json.loads(path.read_text())
        contract = CONTRACTS[row['kind']]

        for field in contract['required']:
            assert field in row

        for check in contract['semantic_checks']:
            assert check(row), f'{path.name} violates contract'

The fixture contract is a poor-man's schema. It does not require a new test framework, and it runs in milliseconds. When you change a fixture, you change a contract, so the diff should be reviewed as carefully as a code diff.

Flaky tests are expensive when they stay in the suite. They are even more expensive when someone disables them with @pytest.mark.skip

and forgets why. A TTL quarantine treats the skip as a temporary decision with a deadline.

from datetime import date
import json
from pathlib import Path

QUARANTINE_FILE = Path('tests/ladder/flaky_quarantine.json')

def active_exclusions():
    data = json.loads(QUARANTINE_FILE.read_text())
    today = date.today().isoformat()
    return [
        f'--deselect=tests/ladder/test_properties.py::{name}'
        for name, deadline in data.get('quarantine_until', {}).items()
        if deadline >= today
    ]

if __name__ == '__main__':
    for flag in active_exclusions():
        print(flag)

You can store the TTL as a plain Python dict for readability:

quarantine_until = {
    'test_discount_rounding': '2026-09-08',
}

Then run the ladder with:

python -m pytest $(python tests/ladder/quarantine.py)

Once 2026-09-08

passes, active_exclusions()

stops producing the --deselect

flag, and the test runs again. If it still fails, you fix the root cause. You do not extend the deadline by editing the dict unless you can explain what changed in the test environment.

Rung Failure symptom Action
Property check Hypothesis finds a counterexample Fix the patch, or fix the property if the patch intentionally changes the invariant
Fixture contract Schema assertion fails before domain tests Review the fixture diff as a contract change; do not silently update the test
TTL quarantine Test runs again and fails Treat it as a normal failure; adding a second quarantine needs a written reason

You can prototype this ladder without a dedicated CI machine. MonkeyCode's free model access can draft the initial CONTRACTS

dict and the quarantine helper; its free server option can run the ladder without spending local resources. The script itself stays plain Python and plain pytest. The grading oracle is the test suite, not the service that drafted it.

Property checks do not prove correctness. They search for counterexamples, and the search is bounded by time, seed, and strategy.

Fixture contracts can give false confidence. If production data grows a new field that the fixtures never see, the contract validates your fixture set, not production.

A TTL quarantine is still a skip. Teams that reset the date every week turn the quarantine into a permanent mute button.

Skip this ladder if you have no flaky tests, a stable oracle, and a review process that already rejects untested patches. In that case, the extra scaffolding is maintenance cost without safety gain. Adopt the rungs one at a time; a one-command ladder is easier to defend after each rung has caught one real regression. The cheapest first step is to add an expiration date to your existing skip list before the next merge.

── more in #developer-tools 4 stories · sorted by recency
── more on @monkeycode 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/one-command-to-verif…] indexed:0 read:5min 2026-09-01 ·