13:02
2026-09-14
digline.dev
ai-research
Bad evals, my own: five exercises from two LLM judges
An engineer running two small LLM judges β "brief" for AI feeds and "scout" for Reddit threads β documented five evaluation exercises showing that three runs of the same 21-case brief suite on Septembβ¦