Evidence-Driven Development: Give Your Coding Agent Something to Prove A developer published a tutorial and companion repository demonstrating "evidence-driven development," a workflow for verifying AI coding agents' claims through inspectable experiments rather than trusting their summaries. The project, Tiny Tasks, is a Python and SQLite task list built with only the standard library that ships with two deliberately broken versions and a script, prove.py, that records which promises each version keeps. The author credits Jan Bosch's 2017 book and the EDDOps paper for related prior work on connecting claims, hypotheses and experiments. An AI coding agent finishes the feature. The tests pass. The explanation sounds reasonable. Then someone asks a small question: “What happens if I run that request again?” That question can change the shape of a project. Now “done” needs to mean something observable. The answer needs an experiment, and the experiment needs a way to be wrong. This tutorial builds a tiny task list around that idea. By the end, you will have a working application, two deliberately broken versions, and a folder of evidence showing which promises each version keeps. I use evidence-driven development to describe this working habit: write down a claim, decide what would disprove it, run the relevant experiment, and let the result govern the next decision. The coding agent can help at every step. The evidence must remain inspectable without trusting the agent's summary. The name predates coding agents. Jan Bosch's 2017 book, Speed, Data, and Ecosystems , includes an Evidence-Driven Development chapter. Its contents connect requirements, hypotheses and experiments. I'm using the term for a practical agent-assisted workflow, without claiming to have invented a methodology. Publisher's catalog https://www.routledge.com/Speed-Data-and-Ecosystems-Succeeding-in-a-Software-Centric-World/Bosch/p/book/9781315270685 . There is related work in AI, too. Xia and colleagues describe evaluation-driven development and operations, or EDDOps, as a continuing feedback loop for LLM agents. Their paper addresses a broader agent lifecycle than the small application we will build here. EDDOps paper https://arxiv.org/abs/2411.13768 . TDD still gives us a useful red–green–refactor loop. BDD helps express behavior through examples people can discuss. In this tutorial, EDD names the discipline of keeping the claim, execution conditions, observed result and resulting decision connected. Those practices fit together. TDD https://agilealliance.org/glossary/tdd/ , Given–When–Then https://agilealliance.org/glossary/given-when-then/ . Our app is Tiny Tasks . It adds tasks and lists them. Python and SQLite are enough; the application and experiment runner use only the Python standard library. Use Python 3.10 or newer with SQLite support. The companion repository https://github.com/copyleftdev/evidence-driven-development contains the complete source, cover artwork and recorded evidence. Clone it to follow along: A standalone companion project for Evidence-Driven Development: Give Your Coding Agent Something to Prove . A tiny local task list teaches explicit claims, negative controls fresh-process experiments, nonzero exercise gates and reproducible evidence. Requires Python 3.10+ with its standard SQLite module. No application dependencies, API keys, model account, network service or cloud resources are required. git clone https://github.com/copyleftdev/evidence-driven-development.git cd evidence-driven-development python3 -m tiny tasks add "Water plants" --request-id plant-1 python3 -m tiny tasks add "Water plants" --request-id plant-1 python3 -m tiny tasks list python3 -m unittest discover -s tests -v python3 scripts/prove.py The second create returns the original task with created: false . Use a new request ID for a genuinely new action. Titles are compared after trimming surrounding whitespace Default data: .local/tasks.sqlite ; set a different file using --db PATH before the add or list subcommand. Conflicting request reuse and invalid input return… git clone https://github.com/copyleftdev/evidence-driven-development.git cd evidence-driven-development git checkout 1f03111db02d359b24602992eddae60122d49d74 python3 -m tiny tasks add "Water plants" --request-id plant-1 python3 -m tiny tasks list The checkout selects the source revision used in this article. The add command returns a task. The list command starts a new process and reads it back. The default database lives at .local/tasks.sqlite . The request ID belongs to the user's intended action. When the caller retries that same action, it sends the same ID. A genuinely new task gets a new ID—even if the title happens to match. That small distinction gives us something worth proving. “Make the task list reliable” leaves too much room for interpretation. Here is the experiment contract we wrote before implementing the application: | Claim | A result that disproves it | |---|---| | Acknowledged tasks survive a process ending | A new process cannot read the saved task | | Repeating a create request returns the original task | A retry creates another task or returns a different ID | | A request ID cannot silently acquire a different meaning | Reusing it with a different title succeeds or changes saved data | | Invalid input leaves the list alone | An empty title adds a task | The complete contract is embedded here, with a commit-pinned copy https://github.com/copyleftdev/evidence-driven-development/blob/1f03111db02d359b24602992eddae60122d49d74/experiments/001-reliable-tasks/experiment.json in the repository: The repository records those claims in experiments/001-reliable-tasks/experiment.json . It also declares three restart rounds, eight concurrent retry processes, two independent repetitions, and the scenarios that must actually execute. These numbers are a small teaching workload. They are not a statistical reliability estimate or a throughput target. Before asking an agent for code, give it this kind of task: Build a local task list that adds and lists tasks. First propose observable claims and the outcomes that would disprove them. Preserve acknowledged tasks across fresh processes. Bind each create request ID to one normalized title and one task. Exercise concurrent retries and conflicting reuse. Keep the application small. Use local synthetic data. Do not claim properties the experiment does not exercise. Review the claims yourself. An agent that quietly narrows “survives interruption” into “works twice in the same object” can produce a beautiful test for the wrong promise. The project has a few small documents with distinct jobs: AGENTS.md working rules and commands docs/plan.md task progress and handoff docs/decisions.md decisions and their reasons experiments/001-reliable-tasks/ experiment.json claims written before code README.md findings and limitations runs/