A benchmark is only as good as the model you use to grade it
A developer built a pytest harness to compare five language models—local Llama, GPT, DeepSeek, and two Claude variants—on cost, speed, and quality, spending about 21 cents on API calls. The initial le…