cd /news/large-language-models/before-you-call-an-llm-endpoint-nerf… · home › topics › large-language-models › article
[ARTICLE · art-146703] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Before You Call an LLM Endpoint "Nerfed": A Small Statistics Checklist in Python

A developer published a Python statistics checklist for evaluating claims that an LLM endpoint has been degraded, using Wilson confidence intervals, Fisher's exact test, and a 200,000-trial simulation showing that with 10 identical providers at a 60% true pass rate, a best-worst gap of 8 or more passes out of 20 appears about 29.9% of the time by chance alone. The checklist argues that pre-registering which endpoints and hypotheses to compare is essential, since scanning a table of providers and quoting the best and worst manufactures apparent differences.

by read10 min views1 publishedOct 7, 2026

Every few weeks someone posts "provider X is serving a watered-down model" with a handful of screenshots, and every few weeks the replies split into "same here" and "works fine for me." Both camps are usually arguing from data that can't settle the question.

My last post argued that a single output tells you almost nothing. A commenter on it made a sharper point that I want to build on: even with repeated runs, the way you pick what to compare can manufacture a difference. Comparing two pre-chosen providers at 16/20 vs 8/20 is meaningful (Fisher p ≈ 0.02). Scanning a table of ten providers and quoting the best and worst is not, because with ten identical providers a gap that large shows up by luck about 30% of the time. Their suggested fix — compare against a pooled rate or a reference endpoint chosen in advance — is the backbone of this post.

So here's the checklist I now run before believing that an endpoint regressed. Everything assumes the simplest useful setup: one fixed probe (for example, "draw a pelican riding a bicycle as SVG"), a fixed pass/fail rubric, and a count of passes out of N attempts. All numbers below are illustrative — they come from the code shown, not from any real provider. Every snippet runs as-is with scipy and numpy.

A pass rate without an interval is a vibe. For binomial counts, the Wilson interval behaves well at small N:

from scipy.stats import binomtest

for label, k, n in [("last week", 34, 40), ("this week", 25, 40)]:
    ci = binomtest(k, n).proportion_ci(confidence_level=0.95, method="wilson")
    print(f"{label}: {k}/{n} = {k/n:.1%}   95% Wilson CI {ci.low:.1%} to {ci.high:.1%}")
last week: 34/40 = 85.0%   95% Wilson CI 70.9% to 92.9%
this week: 25/40 = 62.5%   95% Wilson CI 47.0% to 75.8%

The intervals barely overlap. That alone doesn't settle it (overlapping intervals are not a significance test, and non-overlap isn't required for one), but it tells you how loosely 40 runs pin down each rate. If you only had 10 runs, both intervals would span roughly half the scale.

If you decided before running anything that you would compare "last week" against "this week" on this probe, a 2×2 exact test is the honest tool:

from scipy.stats import fisher_exact

table = [[34,    6],   # last week (illustrative)
         [25,   15]]   # this week (illustrative)

two_sided = fisher_exact(table, alternative="two-sided").pvalue
one_sided = fisher_exact(table, alternative="greater").pvalue  # only if "drop" was pre-registered
print(f"two-sided p = {two_sided:.4f}")
print(f"one-sided p = {one_sided:.4f}")
two-sided p = 0.0406
one-sided p = 0.0203

Two caveats. The one-sided p-value is only legitimate if "this week is worse" was the hypothesis you wrote down in advance; choosing the direction after seeing the data is just halving your p-value for free. And p = 0.04 means "data this lopsided would be unusual if nothing changed," not "96% chance the provider swapped models."

This is the commenter's point, and it deserves a simulation. Give every provider the exact same true pass rate (60%), run each 20 times, and look at the gap between the best and worst:

import numpy as np

rng = np.random.default_rng(42)
n_runs, p_true, sims = 20, 0.6, 200_000   # every provider is identical: 60% pass rate

for k in (2, 5, 10, 20):
    passes = rng.binomial(n_runs, p_true, size=(sims, k))
    gap = passes.max(axis=1) - passes.min(axis=1)
    print(f"{k:>2} providers: P(best-worst gap >= 8) = {(gap >= 8).mean():5.1%}   "
          f"95th percentile of gap = {np.percentile(gap, 95):.0f}")
2 providers: P(best-worst gap >= 8) =  1.4%   95th percentile of gap = 6
 5 providers: P(best-worst gap >= 8) = 10.4%   95th percentile of gap = 8
10 providers: P(best-worst gap >= 8) = 29.9%   95th percentile of gap = 10
20 providers: P(best-worst gap >= 8) = 62.4%   95th percentile of gap = 11

With two pre-chosen providers, an 8-pass gap is rare. With ten, it happens about 30% of the time with nothing wrong — matching the commenter's estimate. With twenty, it's the norm. The 95th-percentile column is the gap you'd need just to clear the noise floor of "best vs worst," and it keeps rising as the table grows.

If your claim comes from eyeballing a leaderboard, the number of comparisons is the number of rows, not two.

Instead of best-vs-worst, ask a question that doesn't depend on which rows happen to land at the extremes: is any endpoint off from the others by more than chance allows for this many endpoints?

import numpy as np
from scipy.stats import chi2_contingency, fisher_exact

passes = {"A": 22, "B": 20, "C": 23, "D": 19, "E": 21, "F": 12, "G": 22, "H": 20}
n = 30

table = np.array([[k, n - k] for k in passes.values()])
chi2, p_omni, dof, _ = chi2_contingency(table)
print(f"omnibus chi2 = {chi2:.2f}, dof = {dof}, p = {p_omni:.4f}")

raw = {}
for name, k in passes.items():
    others_pass = sum(passes.values()) - k
    others_fail = n * (len(passes) - 1) - others_pass
    raw[name] = fisher_exact([[k, n - k], [others_pass, others_fail]]).pvalue

m = len(raw)
running_max = 0.0
for i, (name, p) in enumerate(sorted(raw.items(), key=lambda kv: kv[1])):
    running_max = max(running_max, min(1.0, (m - i) * p))
    print(f"{name}: {passes[name]}/{n}  raw p = {p:.4f}  Holm p = {running_max:.4f}")
omnibus chi2 = 12.35, dof = 7, p = 0.0895
F: 12/30  raw p = 0.0018  Holm p = 0.0144
C: 23/30  raw p = 0.2221  Holm p = 1.0000
A: 22/30  raw p = 0.4178  Holm p = 1.0000
G: 22/30  raw p = 0.4178  Holm p = 1.0000
E: 21/30  raw p = 0.6862  Holm p = 1.0000
D: 19/30  raw p = 0.8367  Holm p = 1.0000
B: 20/30  raw p = 1.0000  Holm p = 1.0000
H: 20/30  raw p = 1.0000  Holm p = 1.0000

Two things worth noticing:

Holm is a reasonable default: it controls the chance of any false flag across the family and is never less powerful than plain Bonferroni. The leave-one-out pool matters too; if F were included in its own baseline, it would drag the baseline toward itself.

"Pre-registration" sounds academic. In practice it means writing down the comparison before you collect data and making it hard to quietly edit later:

import hashlib, json

plan = {
    "question": "Has endpoint X's first-attempt pass rate dropped vs reference R?",
    "endpoints": {"test": "provider-x/model-y", "reference": "official-api/model-y"},
    "probe": "pelican-static-v1",
    "rubric": "pass = all required checks pass; timeouts count as fail",
    "runs_per_endpoint": 80,
    "schedule": "4 blocks, interleaved, random order within block",
    "primary_test": "Fisher exact, one-sided (X < R), alpha = 0.05",
    "stopping_rule": "no interim looks",
}
blob = json.dumps(plan, sort_keys=True).encode()
print(hashlib.sha256(blob).hexdigest())
1e240df48b2c46c2cbd1782eecd2fe87c69f1ef6b469b3a97f3625e803ab884e

Post the hash (in a commit, an issue, a chat message) before the first run, and publish the plan alongside the results. Note what goes in it: a reference endpoint fixed in advance, the sample size, the schedule, how timeouts are counted, the primary test, and the fact that there are no interim looks. If you later add a second probe or a third provider, that's a new plan, not an edit.

The reference endpoint matters most for "did it regress over time?" questions. Re-run the reference in the same window as the endpoint you suspect. If both dropped together, the culprit is more likely your probe, your judge, or something shared upstream — not a secret model swap at one provider.

Twenty runs per side feels like a lot when you're doing it by hand. It isn't, if the drop you're worried about is moderate. Here's a simulated power check for a real 80% → 60% drop:

import numpy as np
from scipy.stats import fisher_exact

rng = np.random.default_rng(7)
p_before, p_after, sims = 0.80, 0.60, 2000   # a real 20-point drop

for n in (20, 40, 80, 120):
    a = rng.binomial(n, p_before, sims)
    b = rng.binomial(n, p_after, sims)
    hits = sum(
        fisher_exact([[x, n - x], [y, n - y]], alternative="greater").pvalue < 0.05
        for x, y in zip(a, b)
    )
    print(f"n = {n:>3} per arm: power ~ {hits / sims:.0%}")
n =  20 per arm: power ~ 30%
n =  40 per arm: power ~ 56%
n =  80 per arm: power ~ 82%
n = 120 per arm: power ~ 95%

At 20 runs per arm you'd miss a genuine 20-point drop about 70% of the time. Around 80 per arm gets you to the conventional 80% power. Two consequences:

The most natural thing in the world is to check the numbers after every batch and post as soon as the gap looks significant. It quietly breaks the test:

import numpy as np
from scipy.stats import fisher_exact

rng = np.random.default_rng(1)
p_true, sims, looks = 0.7, 2000, range(10, 61, 5)   # no real difference at all

false_alarms = 0
for _ in range(sims):
    a = rng.random(60) < p_true
    b = rng.random(60) < p_true
    for n in looks:
        x, y = a[:n].sum(), b[:n].sum()
        if fisher_exact([[x, n - x], [y, n - y]]).pvalue < 0.05:
            false_alarms += 1
            break

fixed = 0
for _ in range(sims):
    x, y = rng.binomial(60, p_true, 2)
    fixed += fisher_exact([[x, 60 - x], [y, 60 - y]]).pvalue < 0.05

print(f"peek every 5 runs, stop at p<0.05: false alarm rate ~ {false_alarms / sims:.1%}")
print(f"single test at n=60:              false alarm rate ~ {fixed / sims:.1%}")
peek every 5 runs, stop at p<0.05: false alarm rate ~ 8.7%
single test at n=60:              false alarm rate ~ 3.4%

Here both arms are identical. Peeking every 5 runs and stopping at the first p < 0.05 more than doubles the false-alarm rate compared with one test at the planned N. (Fisher's test is conservative at these sizes, which is why the fixed-N rate sits below 5%.) If you genuinely need to monitor continuously, use a method designed for it, such as sequential or always-valid tests. Otherwise, fix N and look once.

None of this tells you why an endpoint got worse. Timeouts, truncated outputs, a different default reasoning effort, and an actual model change can all lower a pass rate. Statistics only tells you whether there's a difference worth explaining. Keep timeouts and wrong answers as separate categories in your logs, so that when something does show up you can tell which kind of failure moved.

The statistics above fit in a page. What actually costs time is the bookkeeping around it: same prompt, same parameters, timestamps on every run, a reference endpoint re-tested in the same window, and failures kept instead of dropped.

That bookkeeping is the problem I work on at Folkbench. It's a public leaderboard that compares different services (official APIs and third-party relays) serving the same model, using one published methodology, with the test time shown on each result and missing data left blank rather than guessed. There's also a pelican-on-a-bicycle gallery, which the site itself labels as a fun side-by-side that does not feed into the rankings and should not be read as a capability score (given everything above, I think that is the right call). If you'd rather start from existing results than build the harness yourself, it's at folkbench.com. Apply the same checklist to it: ask what was tested, when, and how many runs are behind each number.

If you run this kind of comparison yourself, I'd like to hear how you choose the reference endpoint. That's the decision I'm least sure about.

Disclosure: I work on Folkbench. Drafted with AI assistance; code and numbers checked by me. All numbers in this post are illustrative outputs of the code shown, not measurements of any real provider.

── more in #large-language-models 4 stories · sorted by recency
── more on @scipy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/before-you-call-an-l…] indexed:0 read:10min 2026-10-07 · —