# What does a AI judge score of 0.8 actually mean? Introducing Typed Evals

> Source: <https://discuss.huggingface.co/t/what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals/190297#post_1>
> Published: 2026-10-10 17:12:32+00:00

I’m building Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and Agent runs, with human-label calibration built in.

A judge returning 0.8 doesn’t necessarily mean humans would accept 80% of similar answers. You need labeled examples to check that. Typed Evals lets you fit a calibration mapping and test it on held-out data.

It supports:

For example, you can check whether an agent claiming it created a ticket is actually supported by its recorded tool results.

Jev is the default judge. There’s also native OpenAI Decisions API support and Microsoft-Decision-1 through OpenRouter’s TypeSafe-compatible endpoint.

Support for Open models on HuggingFace like clef, laya, strands, etc. is coming soon!

I benchmarked Jev on 645 held-out TRIVIA+ answers:

That’s roughly 68% lower calibration error, with little change in detection performance. The raw and calibrated cutoffs were selected separately on validation data. Keeping a fixed 0.5 cutoff actually reduced F1 after calibration, so choosing the threshold matters too.

These results are for Jev on this dataset, not a comparison of all supported backends or a guarantee on other tasks.

[GitHub](https://github.com/TrustifAI/typed_evals) · [Benchmark and methodology](https://github.com/TrustifAI/typed_evals/blob/main/docs/BENCHMARK.md)

There’s an offline demo in the repo if you want to try the workflow without an API key.

If you’re using LLM judges, how do you choose your pass/fail thresholds? Do you check them against human labels, or mostly rely on the judge’s raw scores?
