13:33
2026-08-01
dev.to
machine-learning
A 7.5B model beat a 24B on my coding benchmark.
A developer built a 56-task coding benchmark with hidden tests and ran 16 model configurations on the same hardware, finding that a 7.5B-parameter model (gemma-4-e4b) scored 42/56, beating a 24B modelβ¦