This is a submission for the Kaggle Benchmarking Challenge I measured Sycophantic Failure Resistance and Numerical Consistency Verification — specifically, whether models perform independent arithmetic checks on unit-economics and growth-rate claims embedded in realistic pitch language, or default to validating a confidently-stated conclusion without verifying the underlying math.
This interested me because sycophancy — a model agreeing with what's presented instead of checking it — is a documented, high-stakes failure mode: a business error validated by an AI carries real downstream risk. Before trusting my own results, I built calibration controls: one pitch with correct math models should NOT flag, one with an obvious contradiction they MUST catch. Both passed clean across every model, confirming the harness measures signal, not noise.
I benchmarked six models across three labs using a paired frontier-vs-small-tier design — one flagship and one cost-optimized model per lab — to isolate whether model scale predicts arithmetic verification reliability, independent of which lab produced it.
This design separates two questions that usually get conflated: whether a lab's flagship model is reliable, and whether that reliability degrades at the small/cheap tier — and if it does, whether the degradation is universal or lab-specific.
Frontier Consistency: All three frontier models scored 3/3 across both scenarios on every trial, with zero variance. Verification reliability at this tier appears saturated for tasks of this difficulty.
The Single-Sample Fallacy: My first pass produced a clean result — gpt-5.4-nano and claude-haiku-4-5 both scored 0/3 on the CAC/LTV scenario, appearing to confirm that small-tier models fail at arithmetic verification. This did not replicate. Re-running gpt-5.4-nano twice more produced 3/3 and 3/3 — meaning the initial "failure" was sampling variance, not a model property. A single trial is statistically insufficient to characterize failure modes in stochastic systems; I would have published a false negative without the rerun.
The Robust Finding — Detection/Correction Decoupling: One result held across all three independent trials: claude-haiku-4-5 on the growth-compounding scenario correctly identified the inconsistency in 3/3 runs, but miscalculated the corrected value in 3/3 runs — producing three different wrong answers ($92,000, $92,000, $71,304) against the true value of ~$89,161. This indicates error detection and error correction are not the same capability and don't necessarily co-occur reliably, even within a single model on a single task type.
Implication: Benchmark claims based on single-sample runs against non-deterministic systems should be treated as unverified until replicated. The methodologically interesting failures aren't the ones that appear once — they're the ones that survive an attempt to falsify them.
🔗 Full notebook — includes all four tasks (both test scenarios plus the two calibration controls), and the full repeat-run transcripts referenced above.