More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning
A controlled study of three compact open-weight language models under five billion parameters found that no single model dominates on mathematical reasoning, with Qwen3:4b (Alibaba) most accurate on t…