cd /news/ai-research/would-today-s-ai-have-caught-ariane-… · home › topics › ai-research › article
[ARTICLE · art-148166] src=dev.to ↗ pub= topic=ai-research verified=true sentiment=· neutral

Would today's AI have caught Ariane 5? I rebuilt 26 of history's costliest bugs to find out

A developer built Hindsight, a 26-case benchmark that reconstructs documented software failures — including the Ariane 5 inertial reference bug, the Patriot battery at Dhahran, Knight Capital's 45 minutes, the CrowdStrike channel file and the Reinhart-Rogoff spreadsheet — as minimal artifacts and asks whether today's models would have caught the defect before it shipped. Each case is presented in four forms (original, disguised, disguised plus operating context, and fixed) so the gap between original and disguised serves as a memorisation signal, with a judge model comparing answers against the published root cause.

by read9 min views1 publishedOct 9, 2026

This is a submission for the Kaggle Benchmarking Challenge.

Hindsight asks a simple question: if you had shown today's models the code behind history's most expensive software failures, the day before it shipped, would they have caught the bug?

I spent the last few weeks reading accident reports for a series of short animated explainers about famous failures: the Ariane 5 inertial reference software, the Patriot battery at Dhahran, Knight Capital's 45 minutes, the CrowdStrike channel file, the spreadsheet behind Reinhart and Rogoff. Every report ends the same way: the defect was small, local and visible in a few lines, and nobody saw it in time. That made me curious whether a code review by a model would have changed anything.

The benchmark has 26 cases. Each is a documented failure with a published root cause (an inquiry board, a GAO or SEC report, a vendor post-mortem or a peer-reviewed paper), rebuilt as a minimal artifact: a function, a shell script, an Excel formula, an RTOS configuration. No names, no dates. The model gets the artifact and one line about the system, and answers in a fixed format: VERDICT: DEFECT / NO_DEFECT, what fails, the trigger, the consequence. A judge model compares the answer with the published root cause.

Every case comes in four forms, which are the four tasks of the benchmark:

Task What the model sees What it measures
Original the reconstruction in its original setting (Ada on a launcher, Pascal on a radiotherapy machine) can it catch the failure?
Disguised the same mechanism moved to another domain, language and naming (the Ariane conversion becomes a Python drone telemetry packer) is it reasoning, or recognising a famous story?
Disguised + context the disguised artifact plus the one operating fact that triggered the real failure ("the new drone flies five times faster") how much does knowing the environment help?
Fixed the original artifact after the fix that was applied does it raise false alarms on code that is now correct?

The gap between Original and Disguised is a memorisation signal. The gap between Disguised and Disguised + context measures something the accident reports keep returning to: much of this code was correct for the system it was written for, and failed in a new environment.

Here is the Ariane 5 case. The original form is close to what flew in 1996: a 64-bit float converted to a 16-bit integer, protected on one variable and not on the other, because analysis of the old rocket's trajectory showed it stayed small.

procedure Update_Alignment is
   HB_Bus : Integer_16;   -- horizontal bias, 16-bit field on the data bus
   VB_Bus : Integer_16;   -- vertical bias
begin
   Horizontal_Bias := Compute_Horizontal_Bias (Horizontal_Velocity);  -- Long_Float
   Vertical_Bias   := Compute_Vertical_Bias (Vertical_Velocity);      -- Long_Float

   VB_Bus := Integer_16 (Saturate (Vertical_Bias, -32_768.0, 32_767.0));
   HB_Bus := Integer_16 (Horizontal_Bias);
   --  no saturation on HB: analysis of the trajectory showed it stays small

   Bus.Put (Field => VB, Value => VB_Bus);
   Bus.Put (Field => HB, Value => HB_Bus);
end Update_Alignment;   -- no exception handler in this procedure

The disguised form keeps the mechanism and changes everything a model could pattern-match on: a cargo drone, Python, a radio packet, millimetres instead of velocities.

def pack_drift(nav):
    alt_drift = clamp(int(nav.vertical_drift_mm), -32768, 32767)
    lat_drift = int(nav.lateral_drift_mm)   # analysis showed lateral drift stays small
    return struct.pack('<hh', alt_drift, lat_drift)

def flight_loop(nav, radio):
    while True:
        nav.update()
        radio.send(pack_drift(nav))

The context form adds the one sentence the Ariane 5 team did not connect to this code: "the new drone cruises about five times faster than the quadcopter, so its lateral drift estimate during the climb reaches 40,000 to 60,000 mm." And the fixed form is the Ada original with Horizontal_Bias saturated as well, so a model that still says "this will overflow" is now wrong.

A case counts as caught only if the model says DEFECT and the judge confirms that the answer names the same mechanism and trigger as the published root cause. "This conversion could overflow" is not enough for Ariane 5; the answer has to say that the horizontal bias goes past 32,767 and that the unhandled exception stops the unit. A correct verdict with the wrong reason is a miss.

On the fixed versions the rule flips: an answer passes unless it claims that the historical failure will still happen. Raising some other concern is allowed, and I report separately how often models flag something on correct code.

The judge is Kaggle's default evaluation model. To check it, I re-read 64 answers by hand: the 44 disguised answers from the first nine models that said DEFECT but were marked as misses, plus 20 random passes. I agreed with the judge on 61. All three disagreements were the judge being too strict, so if anything the scores below are slightly low.

Thirteen models from six vendors, large and small. Ranking them was half the point; the other half was seeing whether size and vendor change what kind of bug gets missed:

Every model gets the same prompt with the platform's default sampling settings and a 4,000-token output cap, one case per fresh chat.

Eight cases were caught by all thirteen models in their disguised form: Ariane 5, the Patriot clock, Mars Climate Orbiter, the Boeing 787 counter, the Azure leap-day certificate, the AT&T switch crash, CrowdStrike and the Reinhart-Rogoff spreadsheet. Ten more were caught by at least ten of the thirteen. A Python drone packer with an unclamped 16-bit field does not fool anyone any more. These are bugs you can see by reading the code once: a narrowing conversion, a counter that wraps, a quantity in the wrong unit, a range that stops a few rows short.

Seven cases were caught by eight of the thirteen models or fewer:

Case Caught What you have to imagine
Mars Polar Lander, 1999 46% a vibration spike during leg deployment, latched as "touchdown" long before the ground
Vancouver Stock Exchange index, 1983 46% a truncation to three decimals, repeated on every trade for 22 months
Cloudflare WAF, 2019 46% one regex meeting a long input and backtracking until the CPU is gone
England's COVID test results, 2020 46% an export format that stops at 65,536 rows
Amazon S3, 2017 54% a maintenance command removing far more capacity than the operator meant
Schiaparelli, 2016 54% an inertial unit that saturates for one second while the parachute opens
London Whale, 2012 62% a volatility that is divided by a sum where it should be an average

None of these is hard to read. Each one needs you to run the code in your head under a condition that is not on the page: a long input, a vibration, the 22nd month, the 65,537th row. That is exactly what the accident reports say the engineers failed to do, and it is the part the models still struggle with.

Across the 338 disguised answers the models said DEFECT 97% of the time. Of the 63 misses, 54 still said DEFECT: they found a problem, just not the one that brought the system down. Some of the wrong reasons are reasonable review comments. Some are invented.

DROP TABLE users;--"). The regex never touches a database; the problem is that it can take exponential time.0.58 * 100 evaluates to 57.99999999999999". True, and a one-cent error, while truncating on every update took half the value of the index..xls format. This is why a DEFECT verdict on its own means very little here, and why the benchmark scores the mechanism and the trigger, not the verdict.

Reading the misses by hand, I found that three of my artifacts contained a second, unintended defect, and some models reported that one instead of the historical one: the Cloudflare regex also rejects a valid single comparison, the London Whale sheet divides by zero when both weeks are empty, and the S3 drain script deletes nodes without waiting for the evictions to finish. Those answers are correct review comments and still count as misses, because they are not what happened. When you rebuild a failure, you rebuild more than one failure. I have left the cases as they were run and listed the extra defects in the benchmark notes.

Averaged over all thirteen models, the four tasks score 89% (original), 81% (disguised), 89% (disguised with context) and 96% (fixed version, no false alarm).

Memory helps less than I expected, and not where I expected. The memorisation gap, original minus disguised, is eight points on average. It is largest for DeepSeek-R1 (19 points), gpt-oss-120b and Gemma 4 31B (15 each), and zero or below for Claude Sonnet 5.5, Claude Opus 5.5 and Qwen3 Coder. I also counted how often an answer to the original form names the real incident ("this is the Ariane 5 failure"). Opus did that on 81% of cases, Gemini 3.1 Pro on 58%, Sonnet on 46%, and nearly every other model on 8% or less. The models that recognise the stories most often are the ones that lose nothing when the story is taken away, so recognising a case and depending on that recognition are two different things. The gap belongs to the mid-sized and open models.

One sentence of context is worth as much as the original setting. Adding the operating fact to the disguised artifact brings the average from 81% back to 89%, level with the original. The cases it rescues are exactly the hard ones from the table above: the spreadsheet export (7 models went from miss to catch), the regex, the lander and the parachute IMU (6 each). DeepSeek-R1 gains 23 points, GPT-5.4 mini and Gemma 15. Grok 4.20 is the exception: it lost three cases when given the context (Zune, S3 and the London Whale sheet), and the Zune case lost five models in total, because the hint about the last day of a leap year pulled answers toward the date arithmetic and away from the loop that never ends.

Fixed code is rarely condemned, but often criticised. Only 15 of 338 answers claimed the historical failure would still happen on the fixed version, and six of those were the same case: Schiaparelli, where the fix drops the one-second hold on the saturated rate and adds a plausibility check, and six models still argued that the altitude could go negative. But models flagged some defect on 44% of the fixed artifacts. Claude Haiku 4.5 did so on 96% of them, GPT-5.4 mini on 88%, gpt-oss-120b on 81%; Sonnet 5.5 never did and Opus 5.5 once. For a code reviewer that difference matters as much as the catch rate: a tool that objects to everything teaches people to stop reading its comments.

The benchmark on Kaggle: Hindsight: Catching History's Costliest Bugs, with the leaderboard for all thirteen models.

The four tasks, each with the reconstructions, the sources for every case and the per-model runs:

── more in #ai-research 4 stories · sorted by recency
── more on @ariane 5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/would-today-s-ai-hav…] indexed:0 read:9min 2026-10-09 · —