{"slug": "formally-deriving-programs-from-specifications-bird-meertens-using-lean-4", "title": "Formally deriving programs from specifications (Bird-Meertens) using Lean 4", "summary": "Satnam Singh has replayed Richard Bird's 1989 Bird–Meertens derivation of Kadane's algorithm in Lean 4, publishing the verified code in the Kadane.lean file on GitHub. The derivation is a single calc block, mss_eq_kadane, that steps from an O(n³) specification of the maximum-sum contiguous segment to Kadane's O(n) algorithm, with each of the seven lines citing one law — definition of segs, map promotion, fold promotion, map distributivity, Horner's rule, the scan lemma, and fold–scan fusion — and each line checked by Lean using plain core Lean with no Mathlib. For the input [-2, 1, -3, 4, -1, 2, 1, -5, 4] the maximum segment is [4, -1, 2, 1], summing to 6, and the empty segment counts so the answer is never negative.", "body_md": "# Formally deriving programs from specifications using Lean 4 in the Bird-Meertens style\n\nThis page describes the systematic derivation of an efficient algorithm from an obviously correct but inefficient specification using formally verified transformations, as illustrated in the code below (from [`Kadane.lean`](https://github.com/satnam6502/bird-meertens/blob/main/kadane/Kadane.lean)). With the recent advances in AI coding agents and the automation of proofs using AI theorem provers, this inspiring idea from the 1980s deserves another look.\n\n*Tap or click the image to open it at full resolution.*\n\nThe problem: give me a list of integers and ask for the contiguous segment with the largest\nsum, and the obvious thing to do is to try every segment, add each one up, and\nkeep the biggest. For `[-2, 1, -3, 4, -1, 2, 1, -5, 4]` the winner is\n`[4, -1, 2, 1]`, which sums to 6. This brute force approach is easy to believe\nbut it costs O(n³). Kadane’s algorithm gets the same answer in a single O(n)\npass over the list, but it is not at all obvious why it works. Richard Bird\nshowed how to *calculate* Kadane’s algorithm from the obvious specification,\none algebraic law at a time, in *Algebraic Identities for Program Calculation*\n(The Computer Journal, 1989, §8). It is the poster child of the Bird–Meertens\nformalism (affectionately known as Squiggol), and Wikipedia shows the same\nderivation in its\n[Bird–Meertens formalism](https://en.wikipedia.org/wiki/Bird%E2%80%93Meertens_formalism)\narticle.\n\nHere I replay Bird’s derivation in Lean 4. The specification is at the top, Kadane’s algorithm is at the bottom, and every step in between is a law that Lean has checked (no hand waving, no “it is easy to see that”). The bit I like best is that every line of the derivation is itself a runnable program, so we can time each one and watch the complexity drop as the laws are applied. Everything uses plain core Lean with no Mathlib.\n\n## The Derivation\n\nThe whole derivation is one `calc` block, `mss_eq_kadane` in\n[`Kadane.lean`](https://github.com/satnam6502/bird-meertens/blob/main/kadane/Kadane.lean). Each line is a program and each\nstep cites exactly one law, just like Bird’s figure.\n\n```\n  maxL ∘ map sum ∘ segs                                 O(n³)\n= maxL ∘ map sum ∘ concat ∘ map tails ∘ inits           definition of segs\n= maxL ∘ concat ∘ map (map sum) ∘ map tails ∘ inits     map promotion\n= maxL ∘ map maxL ∘ map (map sum) ∘ map tails ∘ inits   fold promotion\n= maxL ∘ map (maxL ∘ map sum ∘ tails) ∘ inits           map distributivity\n= maxL ∘ map (foldl (⊙) 0) ∘ inits                      Horner's rule     O(n²)\n= maxL ∘ scanl (⊙) 0                                    scan lemma        O(n)\n= fst ∘ foldl (⊗) (0, 0)                                fold–scan fusion  O(n)\n```\n\nThe first line is the specification: take all the segments (every tail of\nevery prefix), sum each one, and take the maximum. The next four steps just\nshuffle the plumbing around without changing the cost. Horner’s rule is where\nthe magic happens: the best sum of a segment ending at some point can be\ncomputed with a left fold of `a ⊙ b = max (a + b) 0`, which saves a factor of\n`n`. The scan lemma notices that folding over every prefix is just a `scanl`,\nwhich saves another factor of `n`. Finally, fold–scan fusion carries the\nrunning maximum along with the fold, so we never build the intermediate list\nat all. The pair operator is `(u, v) ⊗ x = (max u w, w)` where `w = v ⊙ x`,\nand that last line is Kadane’s algorithm.\n\nA few details that matter in the Lean version:\n\n- The lists hold `Int` values. The empty segment counts as a segment, so the\nanswer is never negative (for`[-3, -1, -2]` the answer is 0).\n- `maxL` folds`max` starting from`0` , which makes`0` its unit. This is what\nlets the empty segment and the laws play nicely together.\n- The laws used in the middle of the pipeline carry a trailing `∘ g` so that`rw` can find them inside a longer composition.\n\n## Running Times\n\nThe benchmark in [`bench/KadaneBench.lean`](https://github.com/satnam6502/bird-meertens/blob/main/kadane/bench/KadaneBench.lean) runs every\nline of the `calc` block on random lists of growing length and times each one.\n\nBoth axes are logarithmic, so a program that costs `c·nᵏ` shows up as a\nstraight line with slope `k`. The three complexity classes in Bird’s\nderivation turn up as three slopes, which I find very satisfying to see.\n\n- Steps 1 to 5 sit right on top of each other. Their four laws move things around but leave the cost unchanged.\n- Horner’s rule knocks off a factor of `n` , and the scan lemma knocks off\nanother.\n- Fold–scan fusion is still O(n) but it runs about 7 times faster than the scan, because it never builds the intermediate list.\n- At n = 512 the specification takes 271 ms, while the last line takes 2.4 µs. That is about 110,000 times faster for the same answer.\n\n| Step | Law | Bird’s cost | Measured slope | n = 512 | Largest n | Time there | \n|---|---|---|---|---|---|---|\n| 1 | specification | O(n³) | 3.10 | 271 ms | 724 | 807 ms | \n| 2 | definition of segs | O(n³) | 3.10 | 270 ms | 724 | 807 ms | \n| 3 | map promotion | O(n³) | 3.09 | 268 ms | 724 | 803 ms | \n| 4 | fold promotion | O(n³) | 3.12 | 266 ms | 724 | 799 ms | \n| 5 | map distributivity | O(n³) | 3.13 | 265 ms | 724 | 796 ms | \n| 6 | Horner’s rule | O(n²) | 2.10 | 3.39 ms | 5,793 | 611 ms | \n| 7 | scan lemma | O(n) | 0.99 | 16.5 µs | 131,072 | 4.39 ms | \n| 8 | fold–scan fusion | O(n) | 1.02 | 2.42 µs | 131,072 | 661 µs | \n\nThe slopes of the first six lines come out a little above 3 and 2. My guess\nis that `inits` keeps all `n²/2` prefix cells alive, so the memory traffic\ngrows a bit faster than the operation count.\n\nYou might worry that the benchmark is timing a slightly mistyped copy of a\nline rather than the real thing. It is not: `lines_eq_mss` proves that every\ntimed program equals `mss`, and its proof replays the same laws. Each run also\nchecks that all eight lines agree on every input, just to be sure.\n\nEach point is the fastest of three batches, where a batch repeats the call\nuntil it lasts at least 20 ms. A line is dropped after its first call that\ntakes over 500 ms (which is why the cubic lines give up early). A slope is a\nleast-squares fit of `log t` against `log n`, using only calls of 100 µs or\nmore.\n\nThese numbers come from a single run on a 2.60 GHz Intel Xeon (a GCP VM) with Lean 4.28.0. Your absolute times will be different on another machine, but the slopes should hardly change. To run it yourself, from the root of the repository:\n\n```\nlake exe kadane_bench        # takes about 40 seconds\nlake exe kadane_bench 50     # stop each line at 50 ms instead of 500 ms\n```\n\nThis writes every measurement to\n[`bench/kadane_timings.csv`](https://github.com/satnam6502/bird-meertens/blob/main/kadane/bench/kadane_timings.csv) and redraws both SVG\nplots (a light one and a dark one, picked to match your GitHub theme).\n\n## AI Coding and AI Proofs\n\nThis approach of deriving programs from specifications was a great idea from the 1980s which was perhaps ahead of its time but I think has now found relevance in the age of AI coding. Specifically, AI coding agents and AI theorem provers can now work to synthesize and optimize code from specifications or draft implementations into efficient and correct by construction code (the guarantee comes from the checked proofs, not from the agent).\n\nPreviously the level of skill required and the laborious details needed to perform the proofs for practical programs made this approach difficult to apply. Now AI coding and automatic AI theorem proving advances mean we should look again at this approach for synthesizing code in a manner that still retains some form of comprehension for humans, given by the stepwise refinement steps which act as a kind of explanation of how code has been transformed and synthesized.\n\n## Finding out more and some other related work\n\nThe GitHub repo [Algebra of Programming in Agda: Dependent Types for Relational Program Derivation](https://github.com/scmu/aopa) contains an Agda implementation of a library inspired by the [Algebra of Programming](https://www.amazon.com/Algebra-Programming-Prentice-Hall-International-Computer/dp/013507245X) as described in a book by Richard Bird and Oege de Moor (Oege is now famous for GitHub Copilot and XBOW).\n\n[Jeremy Gibbons](https://www.cs.ox.ac.uk/people/jeremy.gibbons/) has written an article about [The School of Squiggol: A History of the Bird−Meertens Formalism](https://www.cs.ox.ac.uk/publications/publication13852-abstract.html).\n\n## Building\n\nThe toolchain is pinned in [`lean-toolchain`](https://github.com/satnam6502/bird-meertens/blob/main/lean-toolchain) (Lean\nv4.34.1) and nothing outside core Lean is needed. Run these from the root of\nthe repository.\n\n```\nlake build                   # checks the derivation\nlake exe kadane_bench        # the timings and plots\n```\n\nYou can find all the code used for this article at [https://github.com/satnam6502/bird-meertens](https://github.com/satnam6502/bird-meertens).", "url": "https://wpnews.pro/news/formally-deriving-programs-from-specifications-bird-meertens-using-lean-4", "canonical_source": "https://satnam6502.github.io/bird-meertens/", "published_at": "2026-10-05 23:31:05+00:00", "updated_at": "2026-10-05 23:48:46.529038+00:00", "lang": "en", "topics": ["ai-research", "machine-learning"], "entities": ["Lean 4", "Richard Bird", "Satnam Singh", "Bird–Meertens formalism", "Kadane's algorithm", "Kadane.lean", "Algebraic Identities for Program Calculation"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/formally-deriving-programs-from-specifications-bird-meertens-using-lean-4", "markdown": "https://wpnews.pro/news/formally-deriving-programs-from-specifications-bird-meertens-using-lean-4.md", "text": "https://wpnews.pro/news/formally-deriving-programs-from-specifications-bird-meertens-using-lean-4.txt", "jsonld": "https://wpnews.pro/news/formally-deriving-programs-from-specifications-bird-meertens-using-lean-4.jsonld"}}