{"slug": "what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals", "title": "What does a AI judge score of 0.8 actually mean? Introducing Typed Evals", "summary": "TrustifAI released Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and agent runs with human-label calibration built in. Benchmarking its default judge, Jev, on 645 held-out TRIVIA+ answers produced roughly 68% lower calibration error with little change in detection performance, though keeping a fixed 0.5 cutoff reduced F1 after calibration, so the threshold must be chosen separately on validation data. The toolkit supports native OpenAI Decisions API and Microsoft-Decision-1 via OpenRouter's TypeSafe-compatible endpoint, with support for open HuggingFace models such as clef, laya, and strands coming soon.", "body_md": "I’m building Typed Evals, an open-source Python toolkit for evaluating LLM responses, RAG pipelines, and Agent runs, with human-label calibration built in.\n\nA judge returning 0.8 doesn’t necessarily mean humans would accept 80% of similar answers. You need labeled examples to check that. Typed Evals lets you fit a calibration mapping and test it on held-out data.\n\nIt supports:\n\nFor example, you can check whether an agent claiming it created a ticket is actually supported by its recorded tool results.\n\nJev is the default judge. There’s also native OpenAI Decisions API support and Microsoft-Decision-1 through OpenRouter’s TypeSafe-compatible endpoint.\n\nSupport for Open models on HuggingFace like clef, laya, strands, etc. is coming soon!\n\nI benchmarked Jev on 645 held-out TRIVIA+ answers:\n\nThat’s roughly 68% lower calibration error, with little change in detection performance. The raw and calibrated cutoffs were selected separately on validation data. Keeping a fixed 0.5 cutoff actually reduced F1 after calibration, so choosing the threshold matters too.\n\nThese results are for Jev on this dataset, not a comparison of all supported backends or a guarantee on other tasks.\n\n[GitHub](https://github.com/TrustifAI/typed_evals) · [Benchmark and methodology](https://github.com/TrustifAI/typed_evals/blob/main/docs/BENCHMARK.md)\n\nThere’s an offline demo in the repo if you want to try the workflow without an API key.\n\nIf you’re using LLM judges, how do you choose your pass/fail thresholds? Do you check them against human labels, or mostly rely on the judge’s raw scores?", "url": "https://wpnews.pro/news/what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals", "canonical_source": "https://discuss.huggingface.co/t/what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals/190297#post_1", "published_at": "2026-10-10 17:12:32+00:00", "updated_at": "2026-10-10 17:17:51.365412+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-agents", "mlops", "developer-tools"], "entities": ["Typed Evals", "TrustifAI", "Jev", "OpenAI Decisions API", "Microsoft-Decision-1", "OpenRouter", "HuggingFace", "TRIVIA+"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals", "markdown": "https://wpnews.pro/news/what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals.md", "text": "https://wpnews.pro/news/what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals.txt", "jsonld": "https://wpnews.pro/news/what-does-a-ai-judge-score-of-0-8-actually-mean-introducing-typed-evals.jsonld"}}