cd /news/ai-tools/what-does-a-ai-judge-score-of-0-8-ac… · home › topics › ai-tools › article
[ARTICLE · art-148828] src=discuss.huggingface.co ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

What does a AI judge score of 0.8 actually mean? Introducing Typed Evals

TrustifAI released Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and agent runs with human-label calibration built in. Benchmarking its default judge, Jev, on 645 held-out TRIVIA+ answers produced roughly 68% lower calibration error with little change in detection performance, though keeping a fixed 0.5 cutoff reduced F1 after calibration, so the threshold must be chosen separately on validation data. The toolkit supports native OpenAI Decisions API and Microsoft-Decision-1 via OpenRouter's TypeSafe-compatible endpoint, with support for open HuggingFace models such as clef, laya, and strands coming soon.

read1 min views2 publishedOct 10, 2026

I’m building Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and Agent runs, with human-label calibration built in.

A judge returning 0.8 doesn’t necessarily mean humans would accept 80% of similar answers. You need labeled examples to check that. Typed Evals lets you fit a calibration mapping and test it on held-out data.

It supports:

For example, you can check whether an agent claiming it created a ticket is actually supported by its recorded tool results. Jev is the default judge. There’s also native OpenAI Decisions API support and Microsoft-Decision-1 through OpenRouter’s TypeSafe-compatible endpoint.

Support for Open models on HuggingFace like clef, laya, strands, etc. is coming soon!

I benchmarked Jev on 645 held-out TRIVIA+ answers:

That’s roughly 68% lower calibration error, with little change in detection performance. The raw and calibrated cutoffs were selected separately on validation data. Keeping a fixed 0.5 cutoff actually reduced F1 after calibration, so choosing the threshold matters too.

These results are for Jev on this dataset, not a comparison of all supported backends or a guarantee on other tasks.

GitHub · Benchmark and methodology There’s an offline demo in the repo if you want to try the workflow without an API key.

If you’re using LLM judges, how do you choose your pass/fail thresholds? Do you check them against human labels, or mostly rely on the judge’s raw scores?

── more in #ai-tools 4 stories · sorted by recency
── more on @typed evals 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-does-a-ai-judge…] indexed:0 read:1min 2026-10-10 · —