PennyWise Economic Decisions: Benchmarking Frontier LLMs on Financial Reasoning Efficiency A developer built the PennyWise Economic Decisions benchmark to test how efficiently frontier large language models handle multi-variable financial reasoning across eight economic scenarios, scoring both accuracy and token spend against a theoretical minimum. Google's Gemini 3.5 Flash scored a perfect 100.0 with a 1.0 efficiency rating, spending exactly $0.035 and recording $0 total waste, while the benchmark also covered OpenAI GPT-5.5 and GPT-5.4 mini, Anthropic Claude Sonnet 4.6 and Opus 4.6, xAI Grok 4.6, DeepSeek R1, and Qwen3 235B. The full implementation, scoring logs, and leaderboard are published on Kaggle. What I Benchmarked I created the PennyWise Economic Decisions benchmark task to evaluate how effectively modern large language models handle complex, multi-variable financial reasoning and economic scenario classifications. The task exposes models to 8 distinct economic scenarios designed to test practical fiscal decision-making. The metric measures not only accuracy in choosing the correct financial option but also the model's efficiency —calculating token spend versus minimum necessary spend to reveal which models deliver the best ROI Return on Investment for financial automation pipelines. Models Tested I evaluated an expansive, diverse cross-section of frontier and miniature models to see how scale impacts financial precision and budget allocation: - Google Gemini 3.5 Flash & 2.5 Flash To observe speed and lightweight economic parsing capabilities - OpenAI GPT-5.5 & GPT-5.4 mini To test cutting-edge reasoning against optimized smaller variants - Anthropic Claude Sonnet 4.6 & Claude Opus 4.6 To evaluate strict rule-following and structured output constraints - xAI Grok 4.6 - DeepSeek R1 To evaluate how specialized reasoning models parse financial chains of thought - Qwen3 235B To check open-weights performance at scale Findings The results were incredibly revealing, specifically highlighting the prowess of highly optimized modern architectures: - Perfect Scores at Fraction of the Cost: Gemini 3.5 Flash performed flawlessly across all metrics. It handled the financial logic effortlessly, yielding a PennyWise Score of 100.0 and a 1.0 average efficiency rating . - Zero Waste Efficiency: The most surprising insight was the absolute optimization of the token usage. Gemini 3.5 Flash spent exactly $0.035 to finish the entire suite, recording a Total Waste of $0 , matching the exact theoretical minimum necessary spend. - What's Next: For future iterations of this benchmark, I intend to introduce highly adversarial data inputs, such as conflicting market indicators and complex JSON schema constraints, to push the reasoning ceilings of these frontier models even further. My Benchmark You can explore the full implementation details, model scoring logs, and real-time execution leaderboard directly on Kaggle here: 👉 PennyWise Economic Decisions Task Page https://www.kaggle.com/benchmarks/tasks/reismesquita/pennywise-economic-decisions Note: This entire benchmark suite was successfully built, configured, and executed using a cloud-based Kaggle Notebook directly inside a Chrome browser tab on an iPad