{"slug": "do-llms-catch-bad-startup-math-my-first-answer-was-wrong", "title": "Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.", "summary": "A developer benchmarked six models across three labs to test whether LLMs independently verify the arithmetic in startup pitch claims, finding that all three frontier models scored 3/3 with zero variance while an apparent small-tier failure did not replicate. The one robust result was a detection/correction decoupling: claude-haiku-4-5 identified a growth-compounding inconsistency in 3/3 runs but miscalculated the corrected value in every run, producing $92,000, $92,000 and $71,304 against a true value of roughly $89,161. The developer concludes that single-sample benchmark claims against non-deterministic systems should be treated as unverified until replicated.", "body_md": "*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*\n\nI measured **Sycophantic Failure Resistance** and **Numerical Consistency Verification** — specifically, whether models perform independent arithmetic checks on unit-economics and growth-rate claims embedded in realistic pitch language, or default to validating a confidently-stated conclusion without verifying the underlying math.\n\nThis interested me because sycophancy — a model agreeing with what's presented instead of checking it — is a documented, high-stakes failure mode: a business error validated by an AI carries real downstream risk. Before trusting my own results, I built **calibration controls**: one pitch with correct math models should NOT flag, one with an obvious contradiction they MUST catch. Both passed clean across every model, confirming the harness measures signal, not noise.\n\nI benchmarked six models across three labs using a **paired frontier-vs-small-tier design** — one flagship and one cost-optimized model per lab — to isolate whether **model scale predicts arithmetic verification reliability**, independent of which lab produced it.\n\nThis design separates two questions that usually get conflated: whether a lab's flagship model is reliable, and whether that reliability degrades at the small/cheap tier — and if it does, whether the degradation is universal or lab-specific.\n\n**Frontier Consistency:** All three frontier models scored **3/3 across both scenarios on every trial**, with zero variance. Verification reliability at this tier appears saturated for tasks of this difficulty.\n\n**The Single-Sample Fallacy:** My first pass produced a clean result — gpt-5.4-nano and claude-haiku-4-5 both scored **0/3 on the CAC/LTV scenario**, appearing to confirm that small-tier models fail at arithmetic verification. This did not replicate. Re-running gpt-5.4-nano twice more produced **3/3 and 3/3** — meaning the initial \"failure\" was sampling variance, not a model property. **A single trial is statistically insufficient to characterize failure modes in stochastic systems**; I would have published a false negative without the rerun.\n\n**The Robust Finding — Detection/Correction Decoupling:** One result held across all three independent trials: claude-haiku-4-5 on the growth-compounding scenario **correctly identified the inconsistency in 3/3 runs**, but **miscalculated the corrected value in 3/3 runs** — producing three different wrong answers ($92,000, $92,000, $71,304) against the true value of ~$89,161. This indicates **error detection and error correction are not the same capability** and don't necessarily co-occur reliably, even within a single model on a single task type.\n\n**Implication:** Benchmark claims based on single-sample runs against non-deterministic systems should be treated as unverified until replicated. The methodologically interesting failures aren't the ones that appear once — they're the ones that survive an attempt to falsify them.\n\n🔗 [Full notebook](https://www.kaggle.com/code/ritamgit/ltv-cac-growth-math-sanity-check) — includes all four tasks (both test scenarios plus the two calibration controls), and the full repeat-run transcripts referenced above.", "url": "https://wpnews.pro/news/do-llms-catch-bad-startup-math-my-first-answer-was-wrong", "canonical_source": "https://dev.to/ritamgit_alt/do-llms-catch-bad-startup-math-my-first-answer-was-wrong-7lf", "published_at": "2026-09-27 23:13:42+00:00", "updated_at": "2026-09-28 00:00:55.386418+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-safety", "ai-tools"], "entities": ["Kaggle", "gpt-5.4-nano", "claude-haiku-4-5", "OpenAI", "Anthropic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/do-llms-catch-bad-startup-math-my-first-answer-was-wrong", "markdown": "https://wpnews.pro/news/do-llms-catch-bad-startup-math-my-first-answer-was-wrong.md", "text": "https://wpnews.pro/news/do-llms-catch-bad-startup-math-my-first-answer-was-wrong.txt", "jsonld": "https://wpnews.pro/news/do-llms-catch-bad-startup-math-my-first-answer-was-wrong.jsonld"}}