{"slug": "evidence-driven-development-give-your-coding-agent-something-to-prove", "title": "Evidence-Driven Development: Give Your Coding Agent Something to Prove", "summary": "A developer published a tutorial and companion repository demonstrating \"evidence-driven development,\" a workflow for verifying AI coding agents' claims through inspectable experiments rather than trusting their summaries. The project, Tiny Tasks, is a Python and SQLite task list built with only the standard library that ships with two deliberately broken versions and a script, prove.py, that records which promises each version keeps. The author credits Jan Bosch's 2017 book and the EDDOps paper for related prior work on connecting claims, hypotheses and experiments.", "body_md": "An AI coding agent finishes the feature. The tests pass. The explanation sounds reasonable.\n\nThen someone asks a small question: **“What happens if I run that request again?”**\n\nThat question can change the shape of a project. Now “done” needs to mean something observable. The answer needs an experiment, and the experiment needs a way to be wrong.\n\nThis tutorial builds a tiny task list around that idea. By the end, you will have a working application, two deliberately broken versions, and a folder of evidence showing which promises each version keeps.\n\nI use **evidence-driven development** to describe this working habit: write down a claim, decide what would disprove it, run the relevant experiment, and let the result govern the next decision.\n\nThe coding agent can help at every step. The evidence must remain inspectable without trusting the agent's summary.\n\nThe name predates coding agents. Jan Bosch's 2017 book, *Speed, Data, and Ecosystems*, includes an Evidence-Driven Development chapter. Its contents connect requirements, hypotheses and experiments. I'm using the term for a practical agent-assisted workflow, without claiming to have invented a methodology. [Publisher's catalog](https://www.routledge.com/Speed-Data-and-Ecosystems-Succeeding-in-a-Software-Centric-World/Bosch/p/book/9781315270685).\n\nThere is related work in AI, too. Xia and colleagues describe evaluation-driven development and operations, or EDDOps, as a continuing feedback loop for LLM agents. Their paper addresses a broader agent lifecycle than the small application we will build here. [EDDOps paper](https://arxiv.org/abs/2411.13768).\n\nTDD still gives us a useful red–green–refactor loop. BDD helps express behavior through examples people can discuss. In this tutorial, EDD names the discipline of keeping the claim, execution conditions, observed result and resulting decision connected. Those practices fit together. [TDD](https://agilealliance.org/glossary/tdd/), [Given–When–Then](https://agilealliance.org/glossary/given-when-then/).\n\nOur app is **Tiny Tasks**. It adds tasks and lists them. Python and SQLite are enough; the application and experiment runner use only the Python standard library. Use Python 3.10 or newer with SQLite support.\n\nThe [companion repository](https://github.com/copyleftdev/evidence-driven-development) contains the complete source, cover artwork and recorded evidence. Clone it to follow along:\n\nA standalone companion project for **Evidence-Driven Development: Give Your Coding Agent\nSomething to Prove**. A tiny local task list teaches explicit claims, negative controls\nfresh-process experiments, nonzero exercise gates and reproducible evidence.\n\nRequires Python 3.10+ with its standard SQLite module. No application dependencies, API keys, model account, network service or cloud resources are required.\n\n```\ngit clone https://github.com/copyleftdev/evidence-driven-development.git\ncd evidence-driven-development\npython3 -m tiny_tasks add \"Water plants\" --request-id plant-1\npython3 -m tiny_tasks add \"Water plants\" --request-id plant-1\npython3 -m tiny_tasks list\npython3 -m unittest discover -s tests -v\npython3 scripts/prove.py\n```\n\nThe second create returns the original task with `created: false`. Use a new request ID\nfor a genuinely new action. Titles are compared after trimming surrounding whitespace\nDefault data: `.local/tasks.sqlite`; set a different file using `--db PATH` before the\n`add` or `list` subcommand. Conflicting request reuse and invalid input return…\n\n```\ngit clone https://github.com/copyleftdev/evidence-driven-development.git\ncd evidence-driven-development\ngit checkout 1f03111db02d359b24602992eddae60122d49d74\npython3 -m tiny_tasks add \"Water plants\" --request-id plant-1\npython3 -m tiny_tasks list\n```\n\nThe checkout selects the source revision used in this article. The `add` command returns a task. The `list` command starts a new process and reads it back. The default database lives at `.local/tasks.sqlite`.\n\nThe request ID belongs to the user's intended action. When the caller retries that same action, it sends the same ID. A genuinely new task gets a new ID—even if the title happens to match.\n\nThat small distinction gives us something worth proving.\n\n“Make the task list reliable” leaves too much room for interpretation.\n\nHere is the experiment contract we wrote before implementing the application:\n\n| Claim | A result that disproves it | \n|---|---|\n| Acknowledged tasks survive a process ending | A new process cannot read the saved task | \n| Repeating a create request returns the original task | A retry creates another task or returns a different ID | \n| A request ID cannot silently acquire a different meaning | Reusing it with a different title succeeds or changes saved data | \n| Invalid input leaves the list alone | An empty title adds a task | \n\nThe complete contract is embedded here, with a [commit-pinned copy](https://github.com/copyleftdev/evidence-driven-development/blob/1f03111db02d359b24602992eddae60122d49d74/experiments/001-reliable-tasks/experiment.json) in the repository:\n\nThe repository records those claims in `experiments/001-reliable-tasks/experiment.json`. It also declares three restart rounds, eight concurrent retry processes, two independent repetitions, and the scenarios that must actually execute.\n\nThese numbers are a small teaching workload. They are not a statistical reliability estimate or a throughput target.\n\nBefore asking an agent for code, give it this kind of task:\n\n```\nBuild a local task list that adds and lists tasks.\n\nFirst propose observable claims and the outcomes that would disprove them.\nPreserve acknowledged tasks across fresh processes.\nBind each create request ID to one normalized title and one task.\nExercise concurrent retries and conflicting reuse.\n\nKeep the application small. Use local synthetic data.\nDo not claim properties the experiment does not exercise.\n```\n\nReview the claims yourself. An agent that quietly narrows “survives interruption” into “works twice in the same object” can produce a beautiful test for the wrong promise.\n\nThe project has a few small documents with distinct jobs:\n\n```\nAGENTS.md                         working rules and commands\ndocs/plan.md                      task progress and handoff\ndocs/decisions.md                 decisions and their reasons\nexperiments/001-reliable-tasks/\n  experiment.json                claims written before code\n  README.md                      findings and limitations\n  runs/<run-id>/                  actual outputs and source hashes\n```\n\n`AGENTS.md` tells an agent where these things live. The experiment file defines the claim. The run directory records what happened. The decision document explains what we chose because of it.\n\nYou can start with a short instruction:\n\n```\nRead the experiment contract before changing behavior.\nRun the declared scenarios and keep failed outputs.\nCreate a new run directory; never replace an earlier run.\nReport which claims passed, which failed, and what was not tested.\n```\n\nInstructions need maintenance. If a useful constraint already lives in a canonical document, link to it. A growing pile of repeated instructions can make the next agent's job harder. OpenAI's current guidance similarly favors focused instructions and task-specific context. [Guidance on skills and project instructions](https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra).\n\nTiny Tasks stores each task alongside its request ID. A unique database constraint prevents two stored rows from sharing that ID.\n\nCreation uses one transaction. It looks for the ID, compares the normalized title if the ID exists, and inserts only when it is new. The command returns success after the transaction commits.\n\nThe complete `tiny_tasks/store.py` implementation is embedded below. You can also read the [commit-pinned source](https://github.com/copyleftdev/evidence-driven-development/blob/1f03111db02d359b24602992eddae60122d49d74/tiny_tasks/store.py).\n\n“Normalized” has an explicit meaning here: the app removes leading and trailing whitespace before comparing titles. The `add` method shows the transaction boundary; validation, connection setup and listing are included so you can inspect the whole example.\n\nThat boundary matters because checking in one transaction and inserting later would leave room for another process to intervene. SQLite documents the write-transaction behavior of `BEGIN IMMEDIATE` in its [transaction reference](https://www.sqlite.org/lang_transaction.html). Still, plausible code is only a candidate explanation. We have to run the workload we promised to support.\n\nA test can create a `Store`, add a task and immediately list it. That is useful. It does not establish what happens in a new process.\n\nOur experiment runner invokes the actual CLI in subprocesses. Every command starts fresh. For the restart scenario, it adds three tasks in sequence, checking after each successful create that a separate process can read the expected list.\n\nFor retries, it exercises two different situations:\n\nThe second scenario checks the returned identities, the number of `created: true` responses and the rows actually saved. Exactly one task should exist.\n\nThe ignored response is a controlled stand-in for a caller that is unsure whether an action succeeded. We did not cut a network connection or kill a process during a commit. Keeping that distinction visible makes the result useful.\n\nNow comes my favorite part: establish that the experiment can catch the defect it claims to detect.\n\nThe `lab/` directory contains two intentionally wrong implementations. They use the same command shape as the real app:\n\nThese are negative controls: known-bad examples that the checks must reject for a specific reason. They are teaching fixtures, not bugs we are pretending to have discovered accidentally.\n\nThe volatile version must fail the restart claim. The duplicate-accepting version must fail the concurrent retry claim. A syntax error in a broken version would not establish either property, so inspect the saved responses and failure reasons.\n\nReview caught a weakness in the first version of the gate: it accepted a control's failure without checking its reason. A child-process crash could have masqueraded as a successful bug detection. We added regression tests, required nonzero exercise counts for the controls, and made the gate check the targeted failure reason. The earlier run remains in the repository, with that gate limitation recorded.\n\nThere is another way to get a misleading green result: execute nothing.\n\nOur runner tracks how many times each required scenario ran. This intentionally incomplete report must fail even though every claimed outcome says `true`:\n\n```\nempty = {\n    \"claims\": {name: True for name in required_scenarios},\n    \"exercised\": {name: 0 for name in required_scenarios},\n}\n```\n\nA check that refuses zero activity is an **exercise gate**. It answers a basic question before interpreting success: did we actually attempt the scenarios our conclusion depends on?\n\nFrom the companion project:\n\n```\npython3 -m unittest discover -s tests -v\npython3 scripts/prove.py\n```\n\nThe first command runs nine tests: six for application behavior and three for the harness. The second runs the process-boundary experiment, both negative-control implementations and the empty-exercise control.\n\nIn the recorded run, the results were:\n\n| Implementation | Restart | Retry after ignored response | Concurrent retry | Conflicting reuse | Invalid input | \n|---|---|---|---|---|---|\n| Volatile control | Fail | Fail | Fail | Fail | Pass | \n| Duplicate-accepting control | Pass | Fail | Fail | Fail | Pass | \n| SQLite candidate | Pass | Pass | Pass | Pass | Pass | \n\nEach implementation ran twice. The pass/fail outcomes and exercise counts matched between repetitions. The SQLite candidate exercised three restart rounds, one ignored-response retry, eight concurrent submissions, one conflicting reuse and one invalid-input scenario per repetition.\n\nThe two controls failed where intended. The report with zero exercised scenarios was rejected. Those facts allow the overall experiment gate to pass.\n\nYou will find a new directory under `experiments/001-reliable-tasks/runs/`. It contains:\n\nThe captured run used for this article is `20261010T042729Z-633b0705`. Your run will have its own ID. The project preserves the [original run directory](https://github.com/copyleftdev/evidence-driven-development/tree/1f03111db02d359b24602992eddae60122d49d74/experiments/001-reliable-tasks/runs/20261010T042729Z-633b0705) so you can inspect the evidence behind this table, including its [summary](https://github.com/copyleftdev/evidence-driven-development/blob/1f03111db02d359b24602992eddae60122d49d74/experiments/001-reliable-tasks/runs/20261010T042729Z-633b0705/summary.json) and [source manifest](https://github.com/copyleftdev/evidence-driven-development/blob/1f03111db02d359b24602992eddae60122d49d74/experiments/001-reliable-tasks/runs/20261010T042729Z-633b0705/manifest.json).\n\nThe hashes help identify the source files used in a run. They do not establish who performed it, and someone who can rewrite both the files and their manifest can forge a matching set. This is reproducible local evidence, not a signed attestation system.\n\nWe can now choose SQLite for this local application's declared process and retry behavior.\n\nWe have not tested power failure, a corrupt disk, a network filesystem, large task volumes, security boundaries or sustained throughput. We have not measured how often different coding models produce a correct implementation. Those would be separate questions with different experiments.\n\nThis is also where you resist rounding a result into the outcome you wanted. If a target is missed, keep the result and the original threshold visible. A follow-up experiment can ask a better question; it should have a new identity.\n\nA useful agent handoff looks like this:\n\n```\nImplemented: persistent task creation and payload-bound retries.\nVerified: the recorded process-boundary scenarios and negative controls.\nEvidence: the named run directory and source manifest.\nNot tested: power loss, disk corruption, throughput or model reliability.\nDecision: use this implementation for the documented local scope.\n```\n\nThat is enough for another developer—or another agent—to pick up the work without inheriting an unsupported “everything works.”\n\nAn agent can propose the claims, build the application, write the harness and summarize the run. That is convenient, but all four can share the same mistaken assumption.\n\nGive the claim a review before implementation. Inspect the failure cases. Keep an external observation boundary: a subprocess, an HTTP client, an independently written reference calculation, or a recorded user journey appropriate to the property. Check the actual artifacts when an agent reports success.\n\nA second model may help review the work. Its approval is another judgment to examine. For this tutorial, deterministic process results carry the decision; no model judge is needed.\n\nThe app itself does not use an LLM. That is deliberate. You can practice this development workflow with your preferred coding agent while keeping the example's behavior easy to inspect and its runs inexpensive to repeat.\n\nChoose one promise your users care about. Describe an observation that would prove it wrong. Run a test that crosses the relevant boundary. Feed that test a known-bad case. Save the result with enough context that somebody else can challenge it.\n\nThen ask your coding agent for the evidence directory.", "url": "https://wpnews.pro/news/evidence-driven-development-give-your-coding-agent-something-to-prove", "canonical_source": "https://dev.to/copyleftdev/evidence-driven-development-give-your-coding-agent-something-to-prove-1h4k", "published_at": "2026-10-10 05:28:21+00:00", "updated_at": "2026-10-10 05:30:28.980536+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops"], "entities": ["Tiny Tasks", "Jan Bosch", "Xia", "EDDOps", "Python", "SQLite", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/evidence-driven-development-give-your-coding-agent-something-to-prove", "markdown": "https://wpnews.pro/news/evidence-driven-development-give-your-coding-agent-something-to-prove.md", "text": "https://wpnews.pro/news/evidence-driven-development-give-your-coding-agent-something-to-prove.txt", "jsonld": "https://wpnews.pro/news/evidence-driven-development-give-your-coding-agent-something-to-prove.jsonld"}}