# Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

> Source: <https://dev.to/ritamgit_alt/do-llms-catch-bad-startup-math-my-first-answer-was-wrong-7lf>
> Published: 2026-09-27 23:13:42+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

I measured **Sycophantic Failure Resistance** and **Numerical Consistency Verification** — specifically, whether models perform independent arithmetic checks on unit-economics and growth-rate claims embedded in realistic pitch language, or default to validating a confidently-stated conclusion without verifying the underlying math.

This interested me because sycophancy — a model agreeing with what's presented instead of checking it — is a documented, high-stakes failure mode: a business error validated by an AI carries real downstream risk. Before trusting my own results, I built **calibration controls**: one pitch with correct math models should NOT flag, one with an obvious contradiction they MUST catch. Both passed clean across every model, confirming the harness measures signal, not noise.

I benchmarked six models across three labs using a **paired frontier-vs-small-tier design** — one flagship and one cost-optimized model per lab — to isolate whether **model scale predicts arithmetic verification reliability**, independent of which lab produced it.

This design separates two questions that usually get conflated: whether a lab's flagship model is reliable, and whether that reliability degrades at the small/cheap tier — and if it does, whether the degradation is universal or lab-specific.

**Frontier Consistency:** All three frontier models scored **3/3 across both scenarios on every trial**, with zero variance. Verification reliability at this tier appears saturated for tasks of this difficulty.

**The Single-Sample Fallacy:** My first pass produced a clean result — gpt-5.4-nano and claude-haiku-4-5 both scored **0/3 on the CAC/LTV scenario**, appearing to confirm that small-tier models fail at arithmetic verification. This did not replicate. Re-running gpt-5.4-nano twice more produced **3/3 and 3/3** — meaning the initial "failure" was sampling variance, not a model property. **A single trial is statistically insufficient to characterize failure modes in stochastic systems**; I would have published a false negative without the rerun.

**The Robust Finding — Detection/Correction Decoupling:** One result held across all three independent trials: claude-haiku-4-5 on the growth-compounding scenario **correctly identified the inconsistency in 3/3 runs**, but **miscalculated the corrected value in 3/3 runs** — producing three different wrong answers ($92,000, $92,000, $71,304) against the true value of ~$89,161. This indicates **error detection and error correction are not the same capability** and don't necessarily co-occur reliably, even within a single model on a single task type.

**Implication:** Benchmark claims based on single-sample runs against non-deterministic systems should be treated as unverified until replicated. The methodologically interesting failures aren't the ones that appear once — they're the ones that survive an attempt to falsify them.

🔗 [Full notebook](https://www.kaggle.com/code/ritamgit/ltv-cac-growth-math-sanity-check) — includes all four tasks (both test scenarios plus the two calibration controls), and the full repeat-run transcripts referenced above.
