I’m building Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and Agent runs, with human-label calibration built in.
A judge returning 0.8 doesn’t necessarily mean humans would accept 80% of similar answers. You need labeled examples to check that. Typed Evals lets you fit a calibration mapping and test it on held-out data.
It supports:
For example, you can check whether an agent claiming it created a ticket is actually supported by its recorded tool results. Jev is the default judge. There’s also native OpenAI Decisions API support and Microsoft-Decision-1 through OpenRouter’s TypeSafe-compatible endpoint.
Support for Open models on HuggingFace like clef, laya, strands, etc. is coming soon!
I benchmarked Jev on 645 held-out TRIVIA+ answers:
That’s roughly 68% lower calibration error, with little change in detection performance. The raw and calibrated cutoffs were selected separately on validation data. Keeping a fixed 0.5 cutoff actually reduced F1 after calibration, so choosing the threshold matters too.
These results are for Jev on this dataset, not a comparison of all supported backends or a guarantee on other tasks.
GitHub · Benchmark and methodology There’s an offline demo in the repo if you want to try the workflow without an API key.
If you’re using LLM judges, how do you choose your pass/fail thresholds? Do you check them against human labels, or mostly rely on the judge’s raw scores?