cd /news/large-language-models/more-accurate-or-more-efficient-eval… · home topics large-language-models article
[ARTICLE · art-109780] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

A controlled study of three compact open-weight language models under five billion parameters found that no single model dominates on mathematical reasoning, with Qwen3:4b (Alibaba) most accurate on two datasets and Gemma3:4b (Google) on Calculus I, but Gemma3:4b returning roughly three times more correct answers per watt-hour than Qwen3:4b on every dataset while generating far fewer output tokens. The study, posted on arXiv (2608.22048v1), evaluated Gemma3:4b, Phi3:3.8b (Microsoft), and Qwen3:4b across Grade 8 Math, Calculus I, and Advanced Probability and Statistics, and concluded that accuracy alone is insufficient for selecting a local model.

read1 min views1 publishedAug 25, 2026

arXiv:2608.22048v1 Announce Type: new Abstract: Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphasize accuracy while fewer quantify local runtime and energy, characterize failure modes, or apply paired statistical comparisons under controlled conditions. This paper presents a controlled, documented procedure for evaluating locally hosted LLMs on mathematical reasoning. It combines fixed inference settings, hierarchical answer extraction and verification, explicit failure-mode classification, and per-question resource measurement, and reports accuracy with paired significance tests and effect sizes. We demonstrate it in a preliminary study of three compact open-weight models under five billion parameters, Gemma3:4b (Google), Phi3:3.8b (Microsoft), and Qwen3:4b (Alibaba), across datasets spanning Grade 8 Math, Calculus I, and Advanced Probability and Statistics. All models ran through the same local inference server on one workstation, using a shared prompt template, controlled settings, and a matched question set per dataset. No single model dominates. Qwen3:4b is most accurate on two datasets and Gemma3:4b on Calculus I, yet Gemma3:4b returns roughly three times more correct answers per watt-hour than Qwen3:4b on every dataset while generating far fewer output tokens; Qwen3:4b requires substantially more generation time, energy, and output per question. Phi3:3.8b is substantially less accurate on all three datasets; its low extraction-failure rate indicates incorrect answers rather than unparsed output, though we caveat possible prompt-format effects. These preliminary findings indicate that accuracy alone is an insufficient basis for selecting a local model.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/more-accurate-or-mor…] indexed:0 read:1min 2026-08-25 ·