cd/entity/OpenAI Evals· home entities OpenAI Evals
grep -l @openai evals /news/*.json | wc -l → 2

OpenAI Evals

mentions 2 type Organization feed RSS

// recent coverage 2 mentions

19:27
2026-07-22
dev.to
artificial-intelligence

An LLM judge is a biased instrument, not a measurement

A developer found that an LLM judge gave opposite results for the same eval run on consecutive days due to position bias, one of three systematic biases documented in the 2023 paper "Judging LLM-as-a-…

// co-occurs with top 8 entities
// topics top 6 topics