22:55
2026-09-23
arize.com
ai-research
Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks
A benchmark of 23,325 judgments across the RAGTruth and SummEval human-labeled datasets found that TypeSafe's Jev matched Claude Opus 5 at 87% accuracy on held-out hallucination detection at roughly 1β¦