Building in Steel, part three of three. Regulation does not need to become mathematics. The thing you do to earn a presumption of compliance does.
Listen to this essay (20 min) Narrated by Charlie via ElevenLabs
Most AI regulation currently in force or in draft consists of paragraphs asking for appropriate human oversight, adequate robustness, and reasonable measures proportionate to the risk.
Those paragraphs do one of two things and usually both. They slow adoption, because nobody can tell in advance whether what they built satisfies them. And they fail to bite, because when the regulator arrives the argument is about what the paragraph meant, which is an argument nobody can win and everybody can afford to have.
The instinct that follows is to write the law in mathematics. That instinct is wrong, and understanding exactly why is what produces the version that works.
Where the cost actually goes #
"Vague rules are bad" is a slogan, and slogans do not tell you what to change. So start with the mechanism.
When an obligation is ambiguous, every regulated party faces interpretation risk: the possibility that a reader with authority takes the words differently than you did. Nobody accepts that risk bare. They over-implement, add controls the rule never asked for, commission an opinion from counsel, and above all they wait, because moving before the ambiguity resolves means potentially moving twice.
That margin is the real cost of vague regulation, and it is invisible in a way that flatters the drafter. It appears in no impact assessment; no firm publishes its interpretation-risk budget. What you see instead is a market adopting more slowly than anyone intended, and everyone blaming caution or culture. A vague obligation is not light-touch regulation. It is a tax collected in delay and legal fees, and the drafter never sees the receipt.
Then notice the second half: that tax buys no safety. A rule too imprecise to check is also too imprecise to enforce at scale, so regulators do the only thing available and sample. They investigate whoever made the news. Everyone else is regulated by the possibility of scrutiny, a mechanism that punishes the conscientious and leaves the negligent alone until something goes wrong publicly.
Slow and unsafe. Not a trade-off between the two. Both, from the same cause.
Why "just formalise the law" is the wrong fix #
Here the engineer says: write the rule as a predicate and the ambiguity disappears. The lawyers are right to reject that, and their case deserves stating at full strength.
Legal systems distinguish rules from standards, and Louis Kaplow's analysis of the choice is worth getting right because it is usually stated wrongly. The distinction is not vague language versus precise language. It is when the content of the law is determined: in advance by the drafter, or afterwards by the adjudicator. "Not more than 30 mph" is settled before you drive. "At a speed reasonable in the conditions" is settled after you crash.
That relocates cost rather than removing it. Specifying in advance is expensive because someone must anticipate every case the rule will govern. Specifying afterwards is cheap to legislate and expensive to live under, because nobody knows where they stand until someone rules. Which gives the criterion: rules pay off when the conduct is frequent enough to amortise the cost of anticipating it. Write a detailed rule for something happening twice a decade and you wasted the drafting. Write a standard for something happening ten thousand times a day and you imposed the uncertainty cost ten thousand times a day.
Vagueness in a legal standard is therefore not sloppiness. It is a designed capability, the one that lets a fifty-year-old statute reach a technology invented last year. For AI in particular, where nobody can enumerate the failure modes, throwing it away would be vandalism: a precise rule about the systems of 2026 will be a loophole by 2031, exploited by exactly the people it was aimed at.
But note where AI conformity assessment falls on Kaplow's criterion. It is about to be one of the highest-frequency compliance interactions in the economy, which is precisely where up-front specification cost is worth paying.
So: do not formalise the standard. The people who want to are solving the wrong half.
Formalise the safe harbour instead #
Leave the standard exactly as vague as it needs to be, and make machine-checkable the thing you do to earn a presumption that you have met it.
Keep "appropriate technical and organisational measures". Alongside it, publish a decidable predicate over an evidence bundle. Satisfy the predicate and you are presumed compliant. The presumption is rebuttable: a regulator who thinks you gamed it can still proceed on the underlying standard, and a court can still find that what you did was not appropriate in your circumstances.
Each half then does the job it is good at. The standard keeps the adaptive, anti-gaming, reaches-the-unforeseen-case property that made it vague. The predicate gives the honest operator something they can run before shipping and get a yes or no, today, without a phone call. It is a floor, not a ceiling: nothing stops a firm doing more, and nothing lets a firm that satisfied the letter while defeating the point walk away, because the standard is still behind it.
One thing a hybrid must answer for. If you keep a standard and add a rule, do you not inherit the costs of both? Partly, and the proposal should be judged on it. It still comes out ahead because the two costs fall on different parties in different proportions. The anticipation cost is paid once, centrally, by the regulator writing the predicate. The uncertainty cost is paid repeatedly, by every regulated party, and today all of them pay it in full. Moving most firms from full uncertainty to a checkable floor, leaving residual uncertainty for the minority near the edge or trying to be clever, is a better distribution than either pure form. Not a free lunch. A cheaper one.
This is not a new legal device #
The strongest thing about this proposal is how unoriginal it is, because conditional presumptions earned by meeting specified conditions are ordinary regulatory furniture.
The EU AI Act already contains exactly this shape. Article 40 provides that high-risk AI systems in conformity with harmonised standards cited in the Official Journal are presumed to be in conformity with the requirements those standards cover. The presumption reaches only as far as the standard reaches, and does not stop anyone showing the conditions were not really met. GDPR uses a softer version through codes of conduct and certification. The DMCA's section 512 safe harbour is a conditional immunity earned by satisfying stated requirements and lost by failing them. The machinery exists, is understood, has been litigated, and is load-bearing.
A note on vocabulary, because these words get used interchangeably and they are four different things done by four different parties. Conformity assessment is checking a product against requirements. Certification is a body attesting the check passed. Accreditation is a national authority attesting the certification body is competent and impartial. Presumption of conformity is the legal effect attaching when the requirements were ones a cited harmonised standard covers. This proposal touches exactly one of the four: it says the requirements being checked should be machine-checkable. It does not accredit anyone, does not replace a certification decision, and does not remove the assessor. It changes what the assessor is reading.
Today that slot is filled with documents reading like the statute they hang from. A standard saying an organisation shall maintain appropriate logging has moved the ambiguity down one level and solved nothing. The presumption is only as useful as the checkability of the thing that earns it, and we are currently spending a genuinely good legal device on prose that reintroduces the problem it was meant to remove.
What already exists, and what this owes to it #
None of the components here are mine, and an essay in this area that does not say so is either uninformed or hoping you are.
Rules as code has a working language. Catala, designed by Denis Merigoux and colleagues at Inria, is a programming language built specifically for statutes, and its central design decision is the one this problem demands: its semantics are default logic, so a general rule can be overridden by an exception, which is how legal texts are actually structured and precisely what general-purpose languages get wrong. It targets tax, benefits and pension computation, and the team has been rewriting the French income tax algorithm in it since June 2023, working with the tax administration and the CNAF benefits agency. To be accurate about status: those are pilots and rewrites, and Catala is not established as the official production system. But the hard part, translating statute into executable form without lying about its structure, has a serious existing answer and I am not proposing a replacement.
Defeasible reasoning is the right formal frame and it is well developed. A safe harbour is a defeasible inference: the presumption holds unless defeated. That is Reiter's default logic and the AI-and-law tradition built on it through Sartor, Prakken, Governatori and Lawsky, whose work on encoding statutory conditions honestly includes marking where a statute is genuinely vague rather than quietly rendering it Boolean. Anything built here should be recognisable to that literature rather than reinventing it worse.
"Code is law" is twenty-five years old and its warning is the live one. Lessig's argument was never that code should implement law. It was that code already regulates, invisibly, without the accountability attaching to legislation. That cuts directly at this proposal, and I take it up below rather than pretending it does not apply.
So the contribution, if there is one, is narrow: not a language, not a logic, not the observation that law can be executable. Only that the presumption of conformity is the right place in the legal machinery to attach this, because it is a slot that already exists, already carries legal effect, and is currently being filled with prose.
What it looks like #
The harmonised standard becomes two artefacts instead of one.
A schema for the evidence bundle: what a conformity claim consists of, in a form that can be validated rather than admired. Model identity and version. Declared purpose. Capability envelope. The human-oversight mechanism, as an interface and not an intention. Test results, with the method named. Incident records. Each field typed, provenance-stamped, and carrying a stated freshness requirement.
A predicate over that bundle, published as an executable artefact: the conditions that must hold for the presumption to attach. Every claim carries a source. No control asserted without evidence of the assessment behind it. The evaluation covers the declared purpose. Nothing stale beyond its window.
Three things then become true that are not true now.
A firm can run the predicate before shipping, and know. That is the interpretation margin gone, which is the speed.
A regulator can run it across every filing rather than a sample. That is enforcement at population scale instead of by newspaper, which is the safety.
And when the predicate says yes but the outcome was bad, everyone can see which condition was too weak and fix that condition. A prose standard that fails gives you an argument. A predicate that fails gives you a line number.
Why fast-and-safe is mechanical here, not rhetorical #
"Move fast and safely" is normally a slogan concealing a trade-off. Here it is not, because removing ambiguity acts on both terms through one channel. Speed rises because the interpretation margin disappears: the waiting, the second opinion, the controls added as insurance against a reading that may never come. Safety rises because a checkable condition is enforceable against everyone rather than against whoever attracted attention.
The expected trade-off comes from imagining safety is produced by strictness. It is mostly produced by coverage. A moderately demanding rule that genuinely applies to every operator beats a stringent one that in practice applies to whoever ends up in the news.
Five objections #
Open texture does not go away, it moves down a level. Hart's objection, and the most philosophically serious. Every rule has a core of settled application and a penumbra where it is genuinely uncertain whether it applies, and that penumbra is a property of language meeting an open-ended world, not a drafting defect. Write a predicate over an evidence bundle and you have not escaped it: a field requiring that testing "covers the declared purpose" contains "covers", and "covers" has a penumbra exactly as "appropriate" does.
This does not defeat the proposal, but the honest response is not that the predicate resolves vagueness. It relocates vagueness to a small number of named places and makes them visible. Instead of one open-textured word governing an entire obligation, you get a schema where nine fields are settled and two are contested, and everybody can see which two. A real gain, and smaller than the one usually advertised. Anyone claiming to have formalised open texture has misunderstood the term, which is why the standard has to stay behind the harbour: that is where the penumbra is supposed to live, adjudicated after the fact by people with authority to do it.
Who writes the predicate? Lessig's objection, and the political one. A statute is enacted by a legislature that can be voted out. A harmonised standard is written by a technical committee most citizens could not name, and if that standard becomes an executable artefact carrying legal effect, whoever controls the artefact exercises regulatory power without the accountability that attaches to exercising it in a statute. Code already regulates without that accountability, which was Lessig's point in 1999, and this proposal makes the code more powerful.
The mitigations are real and partial. The predicate is published rather than proprietary, so it can be read and criticised. It is executable, so its actual effect can be tested rather than argued about, which is more scrutiny than prose standards get today. And it binds only a presumption, not liability, so the legislature's standard remains supreme. Better than the status quo. Not democratic legitimacy, and I would rather say so than dress it up.
Gaming. A precise condition can be satisfied while its purpose is defeated. True, and why the standard must remain behind the harbour rather than be replaced by it. The presumption is rebuttable, the regulator can proceed on the standard regardless, and a firm that games the predicate has bought a burden shift and nothing more. Any version of this that abolishes the standard deserves to fail.
Ossification. Predicates go stale exactly as rules do. Version them, date them, sunset them, and have a conformity claim state which version it was checked against. That is ordinary practice for anyone who has shipped software and completely alien to how standards are currently maintained. Staleness is a solved problem in the discipline being borrowed from.
Distribution, which is the serious one. Formalisation costs money, so machine-checkable compliance could entrench incumbents and crush small entrants. This deserves more than reassurance, and the honest answer inverts it. Ambiguity is already regressive: buying interpretive certainty means retaining a top-tier firm, which is exactly the purchase a small company cannot make, so they either over-comply blindly or gamble. A published, executable predicate is the first version of this a two-person company can run for nothing.
That holds only if the state publishes the predicate and the tooling as public goods, rather than licensing them through a standards body that sells PDFs. Paywalled, it becomes the regressive thing the objection fears. A genuine fork in the road, and it will be decided by procurement officials nobody has heard of.
Who this is actually for #
Any jurisdiction can do this, and the first to do it well sets the format everyone else adopts, because conformity evidence is something firms would very much like to produce once and use everywhere.
That is worth more than it sounds. The frontier-model race leaves no durable position at the end of it. A standards position does, and it does not depreciate when a better model ships, because it sits at the boundary where liability lives rather than inside the technology. It also happens to be achievable by countries that will never out-capitalise the frontier labs, requiring legal drafting quality, mathematical depth and institutional patience rather than compute.
What exists today, plainly #
This section is in all of these essays because the field is bad at tense.
The pieces are real. Conformity evidence as schema-valid data, obligations as decidable predicates, and a gate refusing to let a claim assert as proven anything unchecked: I run that daily against my own requirements, and it catches contradictions before implementation.
What does not exist is any of it at the scale of a regulator. No jurisdiction has published a machine-checkable harmonised standard. The predicates I have written cover obligations narrow enough to be decidable, and a great many real obligations are not, or not yet. The hard unsolved part is not the checking. It is eliciting a rule set both faithful to what the standard meant and decidable, and that work is closer to law than to programming.
So this is a proposal, not a report. The legal device is proven and in force. The formal-methods machinery is proven and in production elsewhere. Nobody has put them in the same room, and that is the actual gap.
The narrow ask #
Not that law become mathematics. Law should stay exactly as vague as it needs to be to reach the case nobody imagined, because that vagueness is a feature bought at a known price.
Only this: the moment a regulation says here is what you can do to be presumed compliant, that part should be something a machine can check.
That sentence already exists in the statute books. It is currently being spent on more prose.
Part one, How High Can You Build in Mud?, sets out the mechanism. Part two, The Gun Did Not Win, asks which countries can absorb it.
About the author: Eduardo Aguilar Pelaez is CTO and co-founder at Legal Engine Ltd. He writes on formal methods, AI agents, and the discipline of building systems that survive being walked away from.