#
What I Benchmarked
I created the PennyWise Economic Decisions benchmark task to evaluate how effectively modern large language models handle complex, multi-variable financial reasoning and economic scenario classifications.
The task exposes models to 8 distinct economic scenarios designed to test practical fiscal decision-making. The metric measures not only accuracy in choosing the correct financial option but also the model's efficiency—calculating token spend versus minimum necessary spend to reveal which models deliver the best ROI (Return on Investment) for financial automation pipelines.
#
Models Tested
I evaluated an expansive, diverse cross-section of frontier and miniature models to see how scale impacts financial precision and budget allocation:
Google Gemini 3.5 Flash & 2.5 Flash (To observe speed and lightweight economic parsing capabilities) #
OpenAI GPT-5.5 & GPT-5.4 mini (To test cutting-edge reasoning against optimized smaller variants) #
Anthropic Claude Sonnet 4.6 & Claude Opus 4.6 (To evaluate strict rule-following and structured output constraints)
- xAI Grok 4.6
DeepSeek R1 (To evaluate how specialized reasoning models parse financial chains of thought) #
Qwen3 235B (To check open-weights performance at scale)
#
Findings
The results were incredibly revealing, specifically highlighting the prowess of highly optimized modern architectures:
Perfect Scores at Fraction of the Cost: Gemini 3.5 Flash performed flawlessly across all metrics. It handled the financial logic effortlessly, yielding aPennyWise Score of 100.0 and a1.0 average efficiency rating . #
Zero Waste Efficiency: The most surprising insight was the absolute optimization of the token usage. Gemini 3.5 Flash spent exactly**$0.035** to finish the entire suite, recording aTotal Waste of $0 , matching the exact theoretical minimum necessary spend. #
What's Next: For future iterations of this benchmark, I intend to introduce highly adversarial data inputs, such as conflicting market indicators and complex JSON schema constraints, to push the reasoning ceilings of these frontier models even further.
#
My Benchmark
You can explore the full implementation details, model scoring logs, and real-time execution leaderboard directly on Kaggle here:
👉 PennyWise Economic Decisions Task Page Note: This entire benchmark suite was successfully built, configured, and executed using a cloud-based Kaggle Notebook directly inside a Chrome browser tab on an iPad!