cd /news/large-language-models/do-llms-catch-bad-startup-math-my-fi… · home › topics › large-language-models › article
[ARTICLE · art-140666] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

A developer benchmarked six models across three labs to test whether LLMs independently verify the arithmetic in startup pitch claims, finding that all three frontier models scored 3/3 with zero variance while an apparent small-tier failure did not replicate. The one robust result was a detection/correction decoupling: claude-haiku-4-5 identified a growth-compounding inconsistency in 3/3 runs but miscalculated the corrected value in every run, producing $92,000, $92,000 and $71,304 against a true value of roughly $89,161. The developer concludes that single-sample benchmark claims against non-deterministic systems should be treated as unverified until replicated.

by read2 min views3 publishedSep 27, 2026

This is a submission for the Kaggle Benchmarking Challenge I measured Sycophantic Failure Resistance and Numerical Consistency Verification — specifically, whether models perform independent arithmetic checks on unit-economics and growth-rate claims embedded in realistic pitch language, or default to validating a confidently-stated conclusion without verifying the underlying math.

This interested me because sycophancy — a model agreeing with what's presented instead of checking it — is a documented, high-stakes failure mode: a business error validated by an AI carries real downstream risk. Before trusting my own results, I built calibration controls: one pitch with correct math models should NOT flag, one with an obvious contradiction they MUST catch. Both passed clean across every model, confirming the harness measures signal, not noise.

I benchmarked six models across three labs using a paired frontier-vs-small-tier design — one flagship and one cost-optimized model per lab — to isolate whether model scale predicts arithmetic verification reliability, independent of which lab produced it.

This design separates two questions that usually get conflated: whether a lab's flagship model is reliable, and whether that reliability degrades at the small/cheap tier — and if it does, whether the degradation is universal or lab-specific.

Frontier Consistency: All three frontier models scored 3/3 across both scenarios on every trial, with zero variance. Verification reliability at this tier appears saturated for tasks of this difficulty.

The Single-Sample Fallacy: My first pass produced a clean result — gpt-5.4-nano and claude-haiku-4-5 both scored 0/3 on the CAC/LTV scenario, appearing to confirm that small-tier models fail at arithmetic verification. This did not replicate. Re-running gpt-5.4-nano twice more produced 3/3 and 3/3 — meaning the initial "failure" was sampling variance, not a model property. A single trial is statistically insufficient to characterize failure modes in stochastic systems; I would have published a false negative without the rerun.

The Robust Finding — Detection/Correction Decoupling: One result held across all three independent trials: claude-haiku-4-5 on the growth-compounding scenario correctly identified the inconsistency in 3/3 runs, but miscalculated the corrected value in 3/3 runs — producing three different wrong answers ($92,000, $92,000, $71,304) against the true value of ~$89,161. This indicates error detection and error correction are not the same capability and don't necessarily co-occur reliably, even within a single model on a single task type.

Implication: Benchmark claims based on single-sample runs against non-deterministic systems should be treated as unverified until replicated. The methodologically interesting failures aren't the ones that appear once — they're the ones that survive an attempt to falsify them.

🔗 Full notebook — includes all four tasks (both test scenarios plus the two calibration controls), and the full repeat-run transcripts referenced above.

── more in #large-language-models 4 stories · sorted by recency
── more on @kaggle 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/do-llms-catch-bad-st…] indexed:0 read:2min 2026-09-27 · —