14:32
2026-10-07
dev.to
large-language-models
Small LLM Judges Approved 11% and 41% of Wrong Answers. Then I Fixed My Own Pairwise Test.
A developer built OracleBench, a benchmark that grades two small open LLM judges (Qwen2.5-3B and Qwen2.5-0.5B) against deterministic oracles — the FAR/DFARS clause registry, GSM8K arithmetic, and 25 r…