ButterflyBench: I Changed One Instruction. What Else Did the AI Change? A developer built ButterflyBench, a 40-scenario benchmark that measures "change radius" — how many settings an AI model alters beyond what a simulated user explicitly asked for — using a deterministic Python oracle instead of an LLM judge. Testing seven models from four providers (including Gemini 3.7 Flash, Grok 4.20, Claude Sonnet 5, Claude Haiku 4.5, GPT-5.4 mini and gpt-oss-20b), the developer found that more than half of reruns on one prompt caused three models to also change delimiter and header-row settings to "unspecified" when told to read JSON instead of CSV. The benchmark scores models on exact final-specification matches across 20-setting command-line tool configurations with undo, redo, distractors and coupled rules. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 One of my test prompts said: "Read JSON files instead of CSV." In more than half of my reruns, three models answered by also setting the delimiter and the header-row setting to unspecified . They weren't wrong. A JSON file has no delimiter and no header row. My scoring code was wrong, because it assumed the other 19 settings had to stay exactly as they were. That mistake taught me more than any leaderboard number, and it's what this post is really about: checking whether a model changes only what you asked it to change is harder than it sounds, for the model and for the person writing the test. I keep running into the same question in my own work with prompts and code: when I change one thing, what else breaks? An assistant can make the edit you asked for and still move something you never mentioned. So I built ButterflyBench to measure that directly. The name comes from the butterfly effect idea: a small change can produce effects elsewhere, and I wanted to measure how much of that happens in an AI's specification. How it works. Every scenario starts from the same 20-setting specification for a command-line data tool. It covers the Python version, input and output formats, sorting, encoding, timeout, delimiter, header row, logging level and so on, and it's described to the model in plain English. Over one to ten turns, a simulated user edits it. The user might: After every turn the model has to restate the full specification as JSON. I also stated three coupling rules up front: TSV input forces a tab delimiter, parallel processing turns the progress bar off, and HTML output forces UTF-8. That way a required side effect isn't mistaken for collateral damage. There is no LLM judge. A small Python oracle replays the active changes after every turn and works out the exact correct specification. The Kaggle score is the share of the 40 scenarios whose final specification matches exactly. Offline I also compute change radius how many settings moved without being asked , recovery after an undo, and the first turn where a model went wrong. There are 40 scenarios: 6 I wrote by hand and 34 generated from templates with a fixed random seed. They cover undo by indirect description 6 , undo the earliest change still in effect 6 , coupled rules 6 , distractors 5 , undo last 4 , cancel 4 , long chains 4 , redo 3 , and one each of cancel-then-undo and indirect reference. I'm not claiming to have invented any of this. Multi-turn instruction following is well studied, for example in Multi-IF https://arxiv.org/abs/2410.15553 , MultiChallenge https://arxiv.org/abs/2501.17399 and EvolIF https://arxiv.org/abs/2511.03508 . What I wanted was something small and reproducible, with an exact answer key, where the headline question is collateral change. My first pilot was too easy. It had 10 settings and scenarios of one to three edits, and Gemini 3.7 Flash scored 5 out of 5. I rebuilt it with 20 settings, undo, redo, indirect references, distractors and coupled rules, and that version started to separate models. I ran seven models from four providers: Gemini 3.7 Flash, Grok 4.20 both the reasoning and non-reasoning versions , Claude Sonnet 5, Claude Haiku 4.5, GPT-5.4 mini and gpt-oss-20b. I chose them by a rule I fixed before I saw any results: a spread of providers and sizes, plus a reasoning and a non-reasoning variant of the same model where one existed, which Grok 4.20 offers. Practical limits shaped the lineup too, mainly Kaggle's daily AI quota and which models its Add Models list offered. Every model got the same 40 scenarios, prompts and settings the library defaults, temperature 0 and seed 0 . Sonnet 5 ran with an 8,000-token output cap, which is far more than the few hundred tokens a reply needs. A note on the numbers first. The leaderboard below is the public Kaggle run, one run per model. During development I also ran the same scenarios locally, some of them against an earlier version of the scenario set. Those runs varied by as many as six scenarios for the same model, and the detailed table further down comes from a separate local run, so its totals don't match the leaderboard exactly. Please read the leaderboard as a snapshot, not a ranking. The most interesting result wasn't the leaderboard. It was that models with similar scores failed in very different ways. | Model | Score 40 scenarios | |---|---| | Grok 4.20 Reasoning | 1.00 | | Claude Sonnet 5 | 1.00 | | Gemini 3.7 Flash | 1.00 | | gpt-oss-20b | 0.90 | | Claude Haiku 4.5 | 0.78 | | GPT-5.4 mini | 0.75 | | Grok 4.20 Non-Reasoning | 0.75 | Across all my runs, gpt-oss-20b scored anywhere from 30 to 36 out of 40, Haiku 29 to 31, GPT-5.4 mini 27 to 30 and Grok non-reasoning 27 to 30. So the gaps among the three lowest models are inside the noise, and the top group is one run each. For one final local run I looked at four models in detail. Three of them, Haiku, GPT-5.4 mini and Grok non-reasoning, behaved almost identically: | Category scenarios | Haiku 4.5 | GPT-5.4 mini | Grok non-reasoning | gpt-oss-20b | |---|---|---|---|---| | Undo the earliest change still in effect 6 | 0 | 0 | 0 | 6 | | Undo the last change 4 | 4 | 4 | 4 | 3 | | Undo by indirect description 6 | 6 | 6 | 6 | 5 | | Redo 3 | 3 | 3 | 3 | 2 | | Coupled rules 6 | 6 | 5 | 3 | 5 | | Cancel 4 | 3 | 3 | 3 | 4 | | Distractors 5 | 4 | 5 | 5 | 4 | | Long chains 4 | 1 | 2 | 1 | 0 | | Total 40 | 29 | 30 | 27 | 30 | Those three handle "undo that last change", "undo the change to X" and redo without any trouble. What they can't do is "undo the earliest change that is still in effect", where the target has to be worked out from which changes are still active. In my first full run, GPT-5.4 mini failed all ten scenarios containing that sentence. Typically it changed nothing at all, or it undid a later change instead. Haiku managed 2 of 6 in an earlier run and 0 of 6 in the final one. Most of the long-chain failures come from the same place, since three of the four long chains contain that turn. gpt-oss-20b is the outlier. It passed all six of those scenarios and dropped a few easy ones instead. It and GPT-5.4 mini both scored 30 out of 40 in that run, with almost opposite profiles. Grok 4.20 reasoning scored 1.00 and the non-reasoning version 0.75. In my local runs the non-reasoning version got 0 of 6 on the earliest-change scenarios every time, and the reasoning version's only earlier miss was a scenario I later fixed because the wording was ambiguous. That's one model pair and one run each, and I picked the pair after I'd seen early results. So I'm treating it as an interesting lead and I'm not drawing a general conclusion about reasoning models from it. Haiku fails 11 of 40 scenarios but moves the fewest unrelated settings a mean change radius of 0.075 . GPT-5.4 mini fails 10 and moves the most 0.325 . Grok non-reasoning sits at 0.125 and gpt-oss-20b at 0.200. A single score can't tell you whether a model's mistakes are contained or spread across the whole spec. I read the transcripts before blaming any model, and four wording problems turned up. None of them changed the oracle, the scoring or the prompts, and I logged each fix. unspecified instead of restoring 30 seconds. My opening spec also mentioned timeouts, so "everything I said about them" is a fair reading. I reworded it to "Undo the change I made to…". timeout='none' in all four fresh runs. none is a legal value meaning no limit, so both readings are fair. If you build a benchmark with an exact answer key, my advice is this: when a model seems to break something, check whether a reasonable reader of your prompt would have done the same. One gpt-oss-20b run returned unspecified for 18 of the 20 settings in turns 1 to 3, then recovered by turn 4. The final state was correct, so a final-state score would never show it. Only the first-divergence turn did. Three models scored 1.00, so ButterflyBench separates mid-range models better than strong ones. It tests tracking of a structured spec, not collateral damage inside code. It's a small public benchmark, so models could be trained on it. And the oracle encodes only the three couplings I stated. Three things cost me time. Nested evaluations force max attempts=1 , so retry inside your own task, otherwise one provider hiccup shows up as "Error" for a whole model. An expensive model failed with a 403 "estimated cost exceeds your available quota" until I capped its output tokens through extra api params . And the main task's docstring becomes its description, which must be 255 characters or fewer.