Do LLMs Catch Bad Startup Math? My First Answer Was Wrong. A developer benchmarked six models across three labs to test whether LLMs independently verify the arithmetic in startup pitch claims, finding that all three frontier models scored 3/3 with zero variance while an apparent small-tier failure did not replicate. The one robust result was a detection/correction decoupling: claude-haiku-4-5 identified a growth-compounding inconsistency in 3/3 runs but miscalculated the corrected value in every run, producing $92,000, $92,000 and $71,304 against a true value of roughly $89,161. The developer concludes that single-sample benchmark claims against non-deterministic systems should be treated as unverified until replicated. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 I measured Sycophantic Failure Resistance and Numerical Consistency Verification — specifically, whether models perform independent arithmetic checks on unit-economics and growth-rate claims embedded in realistic pitch language, or default to validating a confidently-stated conclusion without verifying the underlying math. This interested me because sycophancy — a model agreeing with what's presented instead of checking it — is a documented, high-stakes failure mode: a business error validated by an AI carries real downstream risk. Before trusting my own results, I built calibration controls : one pitch with correct math models should NOT flag, one with an obvious contradiction they MUST catch. Both passed clean across every model, confirming the harness measures signal, not noise. I benchmarked six models across three labs using a paired frontier-vs-small-tier design — one flagship and one cost-optimized model per lab — to isolate whether model scale predicts arithmetic verification reliability , independent of which lab produced it. This design separates two questions that usually get conflated: whether a lab's flagship model is reliable, and whether that reliability degrades at the small/cheap tier — and if it does, whether the degradation is universal or lab-specific. Frontier Consistency: All three frontier models scored 3/3 across both scenarios on every trial , with zero variance. Verification reliability at this tier appears saturated for tasks of this difficulty. The Single-Sample Fallacy: My first pass produced a clean result — gpt-5.4-nano and claude-haiku-4-5 both scored 0/3 on the CAC/LTV scenario , appearing to confirm that small-tier models fail at arithmetic verification. This did not replicate. Re-running gpt-5.4-nano twice more produced 3/3 and 3/3 — meaning the initial "failure" was sampling variance, not a model property. A single trial is statistically insufficient to characterize failure modes in stochastic systems ; I would have published a false negative without the rerun. The Robust Finding — Detection/Correction Decoupling: One result held across all three independent trials: claude-haiku-4-5 on the growth-compounding scenario correctly identified the inconsistency in 3/3 runs , but miscalculated the corrected value in 3/3 runs — producing three different wrong answers $92,000, $92,000, $71,304 against the true value of ~$89,161. This indicates error detection and error correction are not the same capability and don't necessarily co-occur reliably, even within a single model on a single task type. Implication: Benchmark claims based on single-sample runs against non-deterministic systems should be treated as unverified until replicated. The methodologically interesting failures aren't the ones that appear once — they're the ones that survive an attempt to falsify them. 🔗 Full notebook https://www.kaggle.com/code/ritamgit/ltv-cac-growth-math-sanity-check — includes all four tasks both test scenarios plus the two calibration controls , and the full repeat-run transcripts referenced above.