23:37
2026-08-21
dev.to
large-language-models
We built a benchmark, then caught it strangling the models it was grading
Fortitude Omnis Group's OmnisBench benchmark initially showed small models performing surprisingly well, but community feedback revealed potential contamination from old benchmarks like HumanEval and โฆ