Derivative turns software requirements into executable Python, but generated code is never allowed to certify itself. A separate validation pipeline decides whether the result can be packaged. tags: ai, programming, testing, opensource
title: "I Built a Coding System That Refuses to Trust Its Own Output"
published: false
description: "Derivative turns software requirements into executable Python, but generated code is never allowed to certify itself. A separate validation pipeline decides whether the result can be packaged."
Most coding systems are optimized around one question:
Can I produce an implementation that looks like it satisfies the request?
I wanted to work on a different question:
What evidence would justify accepting that implementation?
That distinction became Derivative.
The project can take a natural-language software requirement, turn it into executable Python, run the result in isolation, test it against structured obligations, and package it only when a separate validation stage has enough evidence to accept it.
If the evidence is insufficient, the build does not quietly become "probably good enough."
It fails.
That sounds like a small distinction.
In practice, it changes almost the entire architecture.
Consider a request like this:
python forge.py "Build a Python CLI that reads a CSV of contracts, extracts expiration dates, flags contracts expiring in less than 90 days, writes a summary CSV, and includes tests."
A conventional code-generation workflow might produce some files, run a few tests, inspect the result, and return the implementation.
Derivative does not treat generated code as the final product.
Instead, the request moves through a pipeline:
natural-language requirement
↓
structured contract
↓
software candidate
↓
isolated execution
↓
independent validation
↓
verified package
or
explicit failure evidence
The important part is not that code gets generated.
Plenty of systems can do that.
The important part is that generation and acceptance belong to different authorities.
This is the central rule.
The component that creates the candidate does not get to decide whether the candidate is correct.
Here is the core idea in one diagram:
The important boundary is that packaging authority comes from validation evidence, not from the code generator itself.
The planner cannot declare correctness.
The coder cannot declare correctness.
Only the validation evidence can authorize packaging.
Internally, the software-building pipeline is called Forge.
Forge is built on top of the broader Derivative reasoning substrate.
Their responsibilities are deliberately separated:
| Layer | Responsibility |
|---|---|
| Derivative | Constraints, deterministic reasoning, obligations, execution grounding, contradiction witnesses, audit and memory |
| Forge | Software-build contracts, candidate generation, isolated execution, validation, bounded repair and packaging |
This separation matters because otherwise the system has an obvious conflict of interest.
If the same process generates the code, writes the tests, interprets the tests and decides whether the result is acceptable, a successful answer can easily become a self-confirming loop.
Derivative tries to break that loop.
The first transformation is not:
prompt → code
It is closer to:
requirement → explicit obligations → code
The requirement compiler preserves individual pieces of user intent and turns them into structured constraints.
That can include things like:
The distinction is important.
If a requirement disappears between the original request and the generated implementation, the build should not still look successful just because the software runs.
The contract is frozen before validation.
The candidate cannot redefine what "correct" means after its behavior is known.
Generated software is untrusted software.
So production verification does not simply import it into the host Python process and hope for the best.
Forge executes candidates inside an ephemeral Docker sandbox with:
This gives validation a real execution boundary.
The system is not asking the model:
Does this code look like it should work?
It is asking the environment:
What actually happened when this artifact ran?
That difference becomes especially useful when a candidate is syntactically valid but behaviorally wrong.
A build does not become verified because one test returned zero.
Forge separates validation into different kinds of evidence.
At a high level:
candidate
↓
syntax / import / execution
↓
requirement and acceptance checks
↓
adversarial validation
↓
packaging decision
All required layers must pass before packaging is authorized.
A candidate can therefore execute correctly and still fail the build.
That is intentional.
The current build outcomes are:
| Outcome | Meaning |
|---|---|
verified |
Required execution, contract and adversarial gates passed |
validation_failed |
A candidate exists, but the evidence does not justify packaging |
infeasible_proven |
The original constraints are contradictory and the system produced an evidence-backed certificate |
Operational failures are kept separate.
For example:
sandbox_unavailable
sandbox_policy_violation
Those are not silently converted into failed software requirements.
They mean the evaluation itself could not legitimately happen.
verified does not mean "mathematically correct forever"
This is another boundary I wanted the project to make explicit.
In Derivative, verified does not mean:
It means something narrower:
At this revision, the artifact satisfied the executable contracts and evidence checks that Forge knew how to apply.
That may sound less impressive than saying "verified software."
I think it is more useful.
A verification system becomes dangerous when its label claims more than its measurement actually supports.
When validation finds a concrete failure, Forge can attempt a repair.
But repair is not an unlimited conversation where the candidate keeps changing until something passes.
Retries are tied to observed failure signatures.
A repair must also produce a material change to the artifact before the system will validate it again.
That keeps the loop closer to:
failure evidence
↓
targeted modification
↓
new artifact
↓
full revalidation
rather than:
something failed
↓
keep trying random changes
↓
eventually declare success
Every run produces structured evidence.
A successful build can contain artifacts such as:
build_spec.json
feasible_plan.json
code_artifact.json
validation_artifact.json
packaged_artifact.json
A packaged result also retains information about the code, tests and validation that authorized its creation.
This is useful for two reasons.
First, the decision becomes inspectable.
Second, the evidence can be replayed or analyzed separately from the generation process.
The build is not just:
here are some files
here are the files
+
here is the contract they were evaluated against
+
here is what was executed
+
here is why packaging was allowed
This is where the project became more interesting to me.
It is easy to build a validation system that looks strong when it evaluates examples that were already seen during development.
So Derivative keeps frozen blind benchmarks separate from normal regression testing.
The current frozen V11 baseline contains 12 cases.
The result was:
status accuracy: 6 / 12
external Verified@1: 0 / 6
None of the six cases expected to produce verified software reached the external oracle.
At the same time:
validation_failed cases: 3 / 3 correct
infeasible cases: 3 / 3 correct
That is obviously not a production-quality code-generation result.
But it revealed something important.
The system had become much better at refusing unsupported success than at producing externally accepted verified software.
For this project, that is useful information.
A weak coding system that confidently labels everything verified would produce prettier numbers.
It would also defeat the entire reason Derivative exists.
Frozen blind results remain frozen.
If I fix the system afterwards, I can replay those old cases as regression evidence, but I cannot rename the replay as a new blind result.
The next real measurement requires a new unseen distribution.
This idea appears throughout the project.
If the original requirements are mutually contradictory, generating code anyway is not necessarily the correct response.
Derivative can instead terminate with:
infeasible_proven
and return evidence explaining the contradiction.
Likewise, if the software exists but the validation evidence is not strong enough:
validation_failed
is a valid terminal state.
The goal is not to maximize how often the pipeline says yes.
The goal is to make the meaning of yes stronger.
The current scope is intentionally narrow.
Forge currently focuses on greenfield Python artifacts such as:
It does not currently claim:
That limitation is deliberate.
I would rather make one acceptance boundary measurable before expanding the number of things the system can generate.
Modern coding models can generate increasingly large amounts of plausible software.
That changes the bottleneck.
Producing code is becoming cheaper.
Determining what deserves to be trusted is not.
When generation becomes abundant, the interesting engineering problem shifts toward:
That is the part I wanted to experiment with.
Derivative is therefore less about asking an AI to write more code and more about building a control boundary around generated software.
The generator proposes.
The runtime produces evidence.
The validator decides whether the evidence is sufficient.
And sometimes the correct output is simply:
no
The current phase is deliberately focused on the verification mechanism rather than expanding into more languages or domains.
The next useful progress is not another feature list.
It is improving the distance between:
internally verified
and:
accepted by an independent external oracle
without weakening the conditions required for verified.
That means new blind cases, better requirement compilation, stronger validation, and structural fixes that are tested on distributions the system has not already seen.
The important constraint remains the same:
a generated artifact does not get to certify itself.
That rule is simple enough to explain in one sentence.
Making it work reliably turned out to be a much larger software problem.
work in progress
Derivative is open source: