Before You Call an LLM Endpoint "Nerfed": A Small Statistics Checklist in Python A developer published a Python statistics checklist for evaluating claims that an LLM endpoint has been degraded, using Wilson confidence intervals, Fisher's exact test, and a 200,000-trial simulation showing that with 10 identical providers at a 60% true pass rate, a best-worst gap of 8 or more passes out of 20 appears about 29.9% of the time by chance alone. The checklist argues that pre-registering which endpoints and hypotheses to compare is essential, since scanning a table of providers and quoting the best and worst manufactures apparent differences. Every few weeks someone posts "provider X is serving a watered-down model" with a handful of screenshots, and every few weeks the replies split into "same here" and "works fine for me." Both camps are usually arguing from data that can't settle the question. My last post argued that a single output tells you almost nothing. A commenter on it made a sharper point that I want to build on: even with repeated runs, the way you pick what to compare can manufacture a difference. Comparing two pre-chosen providers at 16/20 vs 8/20 is meaningful Fisher p ≈ 0.02 . Scanning a table of ten providers and quoting the best and worst is not, because with ten identical providers a gap that large shows up by luck about 30% of the time. Their suggested fix — compare against a pooled rate or a reference endpoint chosen in advance — is the backbone of this post. So here's the checklist I now run before believing that an endpoint regressed. Everything assumes the simplest useful setup: one fixed probe for example, "draw a pelican riding a bicycle as SVG" , a fixed pass/fail rubric, and a count of passes out of N attempts. All numbers below are illustrative — they come from the code shown, not from any real provider. Every snippet runs as-is with scipy and numpy . A pass rate without an interval is a vibe. For binomial counts, the Wilson interval behaves well at small N: python from scipy.stats import binomtest Illustrative numbers, not real measurements for label, k, n in "last week", 34, 40 , "this week", 25, 40 : ci = binomtest k, n .proportion ci confidence level=0.95, method="wilson" print f"{label}: {k}/{n} = {k/n:.1%} 95% Wilson CI {ci.low:.1%} to {ci.high:.1%}" last week: 34/40 = 85.0% 95% Wilson CI 70.9% to 92.9% this week: 25/40 = 62.5% 95% Wilson CI 47.0% to 75.8% The intervals barely overlap. That alone doesn't settle it overlapping intervals are not a significance test, and non-overlap isn't required for one , but it tells you how loosely 40 runs pin down each rate. If you only had 10 runs, both intervals would span roughly half the scale. If you decided before running anything that you would compare "last week" against "this week" on this probe, a 2×2 exact test is the honest tool: python from scipy.stats import fisher exact pass fail table = 34, 6 , last week illustrative 25, 15 this week illustrative two sided = fisher exact table, alternative="two-sided" .pvalue one sided = fisher exact table, alternative="greater" .pvalue only if "drop" was pre-registered print f"two-sided p = {two sided:.4f}" print f"one-sided p = {one sided:.4f}" two-sided p = 0.0406 one-sided p = 0.0203 Two caveats. The one-sided p-value is only legitimate if "this week is worse" was the hypothesis you wrote down in advance; choosing the direction after seeing the data is just halving your p-value for free. And p = 0.04 means "data this lopsided would be unusual if nothing changed," not "96% chance the provider swapped models." This is the commenter's point, and it deserves a simulation. Give every provider the exact same true pass rate 60% , run each 20 times, and look at the gap between the best and worst: python import numpy as np rng = np.random.default rng 42 n runs, p true, sims = 20, 0.6, 200 000 every provider is identical: 60% pass rate for k in 2, 5, 10, 20 : passes = rng.binomial n runs, p true, size= sims, k gap = passes.max axis=1 - passes.min axis=1 print f"{k: 2} providers: P best-worst gap = 8 = { gap = 8 .mean :5.1%} " f"95th percentile of gap = {np.percentile gap, 95 :.0f}" 2 providers: P best-worst gap = 8 = 1.4% 95th percentile of gap = 6 5 providers: P best-worst gap = 8 = 10.4% 95th percentile of gap = 8 10 providers: P best-worst gap = 8 = 29.9% 95th percentile of gap = 10 20 providers: P best-worst gap = 8 = 62.4% 95th percentile of gap = 11 With two pre-chosen providers, an 8-pass gap is rare. With ten, it happens about 30% of the time with nothing wrong — matching the commenter's estimate. With twenty, it's the norm. The 95th-percentile column is the gap you'd need just to clear the noise floor of "best vs worst," and it keeps rising as the table grows. If your claim comes from eyeballing a leaderboard, the number of comparisons is the number of rows, not two. Instead of best-vs-worst, ask a question that doesn't depend on which rows happen to land at the extremes: is any endpoint off from the others by more than chance allows for this many endpoints? python import numpy as np from scipy.stats import chi2 contingency, fisher exact Illustrative pass counts for 8 endpoints serving the "same" model, 30 runs each passes = {"A": 22, "B": 20, "C": 23, "D": 19, "E": 21, "F": 12, "G": 22, "H": 20} n = 30 table = np.array k, n - k for k in passes.values chi2, p omni, dof, = chi2 contingency table print f"omnibus chi2 = {chi2:.2f}, dof = {dof}, p = {p omni:.4f}" Each endpoint vs the pooled pass rate of all the OTHER endpoints raw = {} for name, k in passes.items : others pass = sum passes.values - k others fail = n len passes - 1 - others pass raw name = fisher exact k, n - k , others pass, others fail .pvalue Holm step-down correction for 8 tests m = len raw running max = 0.0 for i, name, p in enumerate sorted raw.items , key=lambda kv: kv 1 : running max = max running max, min 1.0, m - i p print f"{name}: {passes name }/{n} raw p = {p:.4f} Holm p = {running max:.4f}" omnibus chi2 = 12.35, dof = 7, p = 0.0895 F: 12/30 raw p = 0.0018 Holm p = 0.0144 C: 23/30 raw p = 0.2221 Holm p = 1.0000 A: 22/30 raw p = 0.4178 Holm p = 1.0000 G: 22/30 raw p = 0.4178 Holm p = 1.0000 E: 21/30 raw p = 0.6862 Holm p = 1.0000 D: 19/30 raw p = 0.8367 Holm p = 1.0000 B: 20/30 raw p = 1.0000 Holm p = 1.0000 H: 20/30 raw p = 1.0000 Holm p = 1.0000 Two things worth noticing: Holm is a reasonable default: it controls the chance of any false flag across the family and is never less powerful than plain Bonferroni. The leave-one-out pool matters too; if F were included in its own baseline, it would drag the baseline toward itself. "Pre-registration" sounds academic. In practice it means writing down the comparison before you collect data and making it hard to quietly edit later: python import hashlib, json plan = { "question": "Has endpoint X's first-attempt pass rate dropped vs reference R?", "endpoints": {"test": "provider-x/model-y", "reference": "official-api/model-y"}, "probe": "pelican-static-v1", "rubric": "pass = all required checks pass; timeouts count as fail", "runs per endpoint": 80, "schedule": "4 blocks, interleaved, random order within block", "primary test": "Fisher exact, one-sided X < R , alpha = 0.05", "stopping rule": "no interim looks", } blob = json.dumps plan, sort keys=True .encode print hashlib.sha256 blob .hexdigest 1e240df48b2c46c2cbd1782eecd2fe87c69f1ef6b469b3a97f3625e803ab884e Post the hash in a commit, an issue, a chat message before the first run, and publish the plan alongside the results. Note what goes in it: a reference endpoint fixed in advance, the sample size, the schedule, how timeouts are counted, the primary test, and the fact that there are no interim looks. If you later add a second probe or a third provider, that's a new plan, not an edit. The reference endpoint matters most for "did it regress over time?" questions. Re-run the reference in the same window as the endpoint you suspect. If both dropped together, the culprit is more likely your probe, your judge, or something shared upstream — not a secret model swap at one provider. Twenty runs per side feels like a lot when you're doing it by hand. It isn't, if the drop you're worried about is moderate. Here's a simulated power check for a real 80% → 60% drop: python import numpy as np from scipy.stats import fisher exact rng = np.random.default rng 7 p before, p after, sims = 0.80, 0.60, 2000 a real 20-point drop for n in 20, 40, 80, 120 : a = rng.binomial n, p before, sims b = rng.binomial n, p after, sims hits = sum fisher exact x, n - x , y, n - y , alternative="greater" .pvalue < 0.05 for x, y in zip a, b print f"n = {n: 3} per arm: power ~ {hits / sims:.0%}" n = 20 per arm: power ~ 30% n = 40 per arm: power ~ 56% n = 80 per arm: power ~ 82% n = 120 per arm: power ~ 95% At 20 runs per arm you'd miss a genuine 20-point drop about 70% of the time. Around 80 per arm gets you to the conventional 80% power. Two consequences: The most natural thing in the world is to check the numbers after every batch and post as soon as the gap looks significant. It quietly breaks the test: python import numpy as np from scipy.stats import fisher exact rng = np.random.default rng 1 p true, sims, looks = 0.7, 2000, range 10, 61, 5 no real difference at all false alarms = 0 for in range sims : a = rng.random 60 < p true b = rng.random 60 < p true for n in looks: x, y = a :n .sum , b :n .sum if fisher exact x, n - x , y, n - y .pvalue < 0.05: false alarms += 1 break fixed = 0 for in range sims : x, y = rng.binomial 60, p true, 2 fixed += fisher exact x, 60 - x , y, 60 - y .pvalue < 0.05 print f"peek every 5 runs, stop at p<0.05: false alarm rate ~ {false alarms / sims:.1%}" print f"single test at n=60: false alarm rate ~ {fixed / sims:.1%}" peek every 5 runs, stop at p<0.05: false alarm rate ~ 8.7% single test at n=60: false alarm rate ~ 3.4% Here both arms are identical. Peeking every 5 runs and stopping at the first p < 0.05 more than doubles the false-alarm rate compared with one test at the planned N. Fisher's test is conservative at these sizes, which is why the fixed-N rate sits below 5%. If you genuinely need to monitor continuously, use a method designed for it, such as sequential or always-valid tests. Otherwise, fix N and look once. None of this tells you why an endpoint got worse. Timeouts, truncated outputs, a different default reasoning effort, and an actual model change can all lower a pass rate. Statistics only tells you whether there's a difference worth explaining. Keep timeouts and wrong answers as separate categories in your logs, so that when something does show up you can tell which kind of failure moved. The statistics above fit in a page. What actually costs time is the bookkeeping around it: same prompt, same parameters, timestamps on every run, a reference endpoint re-tested in the same window, and failures kept instead of dropped. That bookkeeping is the problem I work on at Folkbench. It's a public leaderboard that compares different services official APIs and third-party relays serving the same model, using one published methodology, with the test time shown on each result and missing data left blank rather than guessed. There's also a pelican-on-a-bicycle gallery, which the site itself labels as a fun side-by-side that does not feed into the rankings and should not be read as a capability score given everything above, I think that is the right call . If you'd rather start from existing results than build the harness yourself, it's at folkbench.com https://folkbench.com/?utm source=luntan&utm campaign=dev . Apply the same checklist to it: ask what was tested, when, and how many runs are behind each number. If you run this kind of comparison yourself, I'd like to hear how you choose the reference endpoint. That's the decision I'm least sure about. Disclosure: I work on Folkbench. Drafted with AI assistance; code and numbers checked by me. All numbers in this post are illustrative outputs of the code shown, not measurements of any real provider.