cd/entity/SummEvalΒ· homeβ€Ί entitiesβ€Ί SummEval
grep -l @summeval /news/*.json | wc -l β†’ 3

SummEval

mentions 3 type Organization feed RSS

// recent coverage 3 mentions

22:55
2026-09-23
arize.com
ai-research

Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks

A benchmark of 23,325 judgments across the RAGTruth and SummEval human-labeled datasets found that TypeSafe's Jev matched Claude Opus 5 at 87% accuracy on held-out hallucination detection at roughly 1…

03:19
2026-06-28
arxiv.org
large-language-models

Improved LLM as a Judge Techniques

Researchers propose BINEVAL, a framework that decomposes LLM evaluation into atomic binary questions for interpretable, multi-dimensional scoring. The method matches or outperforms strong baselines on…

// co-occurs with top 8 entities
// topics top 5 topics