What does a AI judge score of 0.8 actually mean? Introducing Typed Evals TrustifAI released Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and agent runs with human-label calibration built in. Benchmarking its default judge, Jev, on 645 held-out TRIVIA+ answers produced roughly 68% lower calibration error with little change in detection performance, though keeping a fixed 0.5 cutoff reduced F1 after calibration, so the threshold must be chosen separately on validation data. The toolkit supports native OpenAI Decisions API and Microsoft-Decision-1 via OpenRouter's TypeSafe-compatible endpoint, with support for open HuggingFace models such as clef, laya, and strands coming soon. I’m building Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and Agent runs, with human-label calibration built in. A judge returning 0.8 doesn’t necessarily mean humans would accept 80% of similar answers. You need labeled examples to check that. Typed Evals lets you fit a calibration mapping and test it on held-out data. It supports: For example, you can check whether an agent claiming it created a ticket is actually supported by its recorded tool results. Jev is the default judge. There’s also native OpenAI Decisions API support and Microsoft-Decision-1 through OpenRouter’s TypeSafe-compatible endpoint. Support for Open models on HuggingFace like clef, laya, strands, etc. is coming soon I benchmarked Jev on 645 held-out TRIVIA+ answers: That’s roughly 68% lower calibration error, with little change in detection performance. The raw and calibrated cutoffs were selected separately on validation data. Keeping a fixed 0.5 cutoff actually reduced F1 after calibration, so choosing the threshold matters too. These results are for Jev on this dataset, not a comparison of all supported backends or a guarantee on other tasks. GitHub https://github.com/TrustifAI/typed evals · Benchmark and methodology https://github.com/TrustifAI/typed evals/blob/main/docs/BENCHMARK.md There’s an offline demo in the repo if you want to try the workflow without an API key. If you’re using LLM judges, how do you choose your pass/fail thresholds? Do you check them against human labels, or mostly rely on the judge’s raw scores?