{"slug": "a-coding-system-that-refuses-to-trust-its-own-output", "title": "A Coding System That Refuses to Trust Its Own Output", "summary": "A developer built Derivative, an open-source coding system that converts natural-language software requirements into executable Python but never lets generated code certify itself. The pipeline separates generation from acceptance: a requirement compiler freezes structured obligations before validation, the candidate runs in isolation, and a separate validation stage must produce sufficient evidence before the result can be packaged — otherwise the build fails with explicit failure evidence rather than being treated as \"probably good enough.", "body_md": "*Derivative turns software requirements into executable Python, but generated code is never allowed to certify itself. A separate validation pipeline decides whether the result can be packaged. tags: ai, programming, testing, opensource*\n\ntitle: \"I Built a Coding System That Refuses to Trust Its Own Output\"\n\npublished: false\n\ndescription: \"Derivative turns software requirements into executable Python, but generated code is never allowed to certify itself. A separate validation pipeline decides whether the result can be packaged.\"\n\nMost coding systems are optimized around one question:\n\n**Can I produce an implementation that looks like it satisfies the request?**\n\nI wanted to work on a different question:\n\n**What evidence would justify accepting that implementation?**\n\nThat distinction became **Derivative**.\n\nThe project can take a natural-language software requirement, turn it into executable Python, run the result in isolation, test it against structured obligations, and package it only when a separate validation stage has enough evidence to accept it.\n\nIf the evidence is insufficient, the build does not quietly become \"probably good enough.\"\n\nIt fails.\n\nThat sounds like a small distinction.\n\nIn practice, it changes almost the entire architecture.\n\nConsider a request like this:\n\n```\npython forge.py \"Build a Python CLI that reads a CSV of contracts, extracts expiration dates, flags contracts expiring in less than 90 days, writes a summary CSV, and includes tests.\"\n```\n\nA conventional code-generation workflow might produce some files, run a few tests, inspect the result, and return the implementation.\n\nDerivative does not treat generated code as the final product.\n\nInstead, the request moves through a pipeline:\n\n```\nnatural-language requirement\n        ↓\nstructured contract\n        ↓\nsoftware candidate\n        ↓\nisolated execution\n        ↓\nindependent validation\n        ↓\nverified package\n        or\nexplicit failure evidence\n```\n\nThe important part is not that code gets generated.\n\nPlenty of systems can do that.\n\nThe important part is that **generation and acceptance belong to different authorities**.\n\nThis is the central rule.\n\nThe component that creates the candidate does not get to decide whether the candidate is correct.\n\nHere is the core idea in one diagram:\n\nThe important boundary is that packaging authority comes from validation evidence, not from the code generator itself.\n\nThe planner cannot declare correctness.\n\nThe coder cannot declare correctness.\n\nOnly the validation evidence can authorize packaging.\n\nInternally, the software-building pipeline is called **Forge**.\n\nForge is built on top of the broader **Derivative** reasoning substrate.\n\nTheir responsibilities are deliberately separated:\n\n| Layer | Responsibility | \n|---|---|\n| Derivative | Constraints, deterministic reasoning, obligations, execution grounding, contradiction witnesses, audit and memory | \n| Forge | Software-build contracts, candidate generation, isolated execution, validation, bounded repair and packaging | \n\nThis separation matters because otherwise the system has an obvious conflict of interest.\n\nIf the same process generates the code, writes the tests, interprets the tests and decides whether the result is acceptable, a successful answer can easily become a self-confirming loop.\n\nDerivative tries to break that loop.\n\nThe first transformation is not:\n\n```\nprompt → code\n```\n\nIt is closer to:\n\n```\nrequirement → explicit obligations → code\n```\n\nThe requirement compiler preserves individual pieces of user intent and turns them into structured constraints.\n\nThat can include things like:\n\nThe distinction is important.\n\nIf a requirement disappears between the original request and the generated implementation, the build should not still look successful just because the software runs.\n\nThe contract is frozen before validation.\n\nThe candidate cannot redefine what \"correct\" means after its behavior is known.\n\nGenerated software is untrusted software.\n\nSo production verification does not simply import it into the host Python process and hope for the best.\n\nForge executes candidates inside an ephemeral Docker sandbox with:\n\nThis gives validation a real execution boundary.\n\nThe system is not asking the model:\n\nDoes this code look like it should work?\n\nIt is asking the environment:\n\nWhat actually happened when this artifact ran?\n\nThat difference becomes especially useful when a candidate is syntactically valid but behaviorally wrong.\n\nA build does not become verified because one test returned zero.\n\nForge separates validation into different kinds of evidence.\n\nAt a high level:\n\n```\ncandidate\n   ↓\nsyntax / import / execution\n   ↓\nrequirement and acceptance checks\n   ↓\nadversarial validation\n   ↓\npackaging decision\n```\n\nAll required layers must pass before packaging is authorized.\n\nA candidate can therefore execute correctly and still fail the build.\n\nThat is intentional.\n\nThe current build outcomes are:\n\n| Outcome | Meaning | \n|---|---|\n| `verified` | Required execution, contract and adversarial gates passed | \n| `validation_failed` | A candidate exists, but the evidence does not justify packaging | \n| `infeasible_proven` | The original constraints are contradictory and the system produced an evidence-backed certificate | \n\nOperational failures are kept separate.\n\nFor example:\n\n```\nsandbox_unavailable\nsandbox_policy_violation\n```\n\nThose are not silently converted into failed software requirements.\n\nThey mean the evaluation itself could not legitimately happen.\n\n`verified` does not mean \"mathematically correct forever\"\nThis is another boundary I wanted the project to make explicit.\n\nIn Derivative, `verified` does **not** mean:\n\nIt means something narrower:\n\nAt this revision, the artifact satisfied the executable contracts and evidence checks that Forge knew how to apply.\n\nThat may sound less impressive than saying \"verified software.\"\n\nI think it is more useful.\n\nA verification system becomes dangerous when its label claims more than its measurement actually supports.\n\nWhen validation finds a concrete failure, Forge can attempt a repair.\n\nBut repair is not an unlimited conversation where the candidate keeps changing until something passes.\n\nRetries are tied to observed failure signatures.\n\nA repair must also produce a material change to the artifact before the system will validate it again.\n\nThat keeps the loop closer to:\n\n```\nfailure evidence\n      ↓\ntargeted modification\n      ↓\nnew artifact\n      ↓\nfull revalidation\n```\n\nrather than:\n\n```\nsomething failed\n      ↓\nkeep trying random changes\n      ↓\neventually declare success\n```\n\nEvery run produces structured evidence.\n\nA successful build can contain artifacts such as:\n\n```\nbuild_spec.json\nfeasible_plan.json\ncode_artifact.json\nvalidation_artifact.json\npackaged_artifact.json\n```\n\nA packaged result also retains information about the code, tests and validation that authorized its creation.\n\nThis is useful for two reasons.\n\nFirst, the decision becomes inspectable.\n\nSecond, the evidence can be replayed or analyzed separately from the generation process.\n\nThe build is not just:\n\n```\nhere are some files\nhere are the files\n+\nhere is the contract they were evaluated against\n+\nhere is what was executed\n+\nhere is why packaging was allowed\n```\n\nThis is where the project became more interesting to me.\n\nIt is easy to build a validation system that looks strong when it evaluates examples that were already seen during development.\n\nSo Derivative keeps frozen blind benchmarks separate from normal regression testing.\n\nThe current frozen V11 baseline contains 12 cases.\n\nThe result was:\n\n```\nstatus accuracy:       6 / 12\nexternal Verified@1:   0 / 6\n```\n\nNone of the six cases expected to produce verified software reached the external oracle.\n\nAt the same time:\n\n```\nvalidation_failed cases: 3 / 3 correct\ninfeasible cases:        3 / 3 correct\n```\n\nThat is obviously not a production-quality code-generation result.\n\nBut it revealed something important.\n\nThe system had become much better at **refusing unsupported success** than at producing externally accepted verified software.\n\nFor this project, that is useful information.\n\nA weak coding system that confidently labels everything `verified` would produce prettier numbers.\n\nIt would also defeat the entire reason Derivative exists.\n\nFrozen blind results remain frozen.\n\nIf I fix the system afterwards, I can replay those old cases as regression evidence, but I cannot rename the replay as a new blind result.\n\nThe next real measurement requires a new unseen distribution.\n\nThis idea appears throughout the project.\n\nIf the original requirements are mutually contradictory, generating code anyway is not necessarily the correct response.\n\nDerivative can instead terminate with:\n\n```\ninfeasible_proven\n```\n\nand return evidence explaining the contradiction.\n\nLikewise, if the software exists but the validation evidence is not strong enough:\n\n```\nvalidation_failed\n```\n\nis a valid terminal state.\n\nThe goal is not to maximize how often the pipeline says yes.\n\nThe goal is to make the meaning of yes stronger.\n\nThe current scope is intentionally narrow.\n\nForge currently focuses on greenfield Python artifacts such as:\n\nIt does not currently claim:\n\nThat limitation is deliberate.\n\nI would rather make one acceptance boundary measurable before expanding the number of things the system can generate.\n\nModern coding models can generate increasingly large amounts of plausible software.\n\nThat changes the bottleneck.\n\nProducing code is becoming cheaper.\n\nDetermining what deserves to be trusted is not.\n\nWhen generation becomes abundant, the interesting engineering problem shifts toward:\n\nThat is the part I wanted to experiment with.\n\nDerivative is therefore less about asking an AI to write more code and more about building a **control boundary around generated software**.\n\nThe generator proposes.\n\nThe runtime produces evidence.\n\nThe validator decides whether the evidence is sufficient.\n\nAnd sometimes the correct output is simply:\n\n```\nno\n```\n\nThe current phase is deliberately focused on the verification mechanism rather than expanding into more languages or domains.\n\nThe next useful progress is not another feature list.\n\nIt is improving the distance between:\n\n```\ninternally verified\n```\n\nand:\n\n```\naccepted by an independent external oracle\n```\n\nwithout weakening the conditions required for `verified`.\n\nThat means new blind cases, better requirement compilation, stronger validation, and structural fixes that are tested on distributions the system has not already seen.\n\nThe important constraint remains the same:\n\n**a generated artifact does not get to certify itself.**\n\nThat rule is simple enough to explain in one sentence.\n\nMaking it work reliably turned out to be a much larger software problem.\n\n*work in progress*\n\nDerivative is open source:", "url": "https://wpnews.pro/news/a-coding-system-that-refuses-to-trust-its-own-output", "canonical_source": "https://dev.to/danielecangi/a-coding-system-that-refuses-to-trust-its-own-output-8dj", "published_at": "2026-10-07 13:35:57+00:00", "updated_at": "2026-10-07 13:47:28.150064+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools", "artificial-intelligence"], "entities": ["Derivative", "Forge"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-coding-system-that-refuses-to-trust-its-own-output", "markdown": "https://wpnews.pro/news/a-coding-system-that-refuses-to-trust-its-own-output.md", "text": "https://wpnews.pro/news/a-coding-system-that-refuses-to-trust-its-own-output.txt", "jsonld": "https://wpnews.pro/news/a-coding-system-that-refuses-to-trust-its-own-output.jsonld"}}