cd/entity/IFBenchΒ· homeβ€Ί entitiesβ€Ί IFBench
grep -l @ifbench /news/*.json | wc -l β†’ 3

IFBench

mentions 3 type Organization feed RSS

// recent coverage 3 mentions

03:19
2026-06-28
arxiv.org
large-language-models

Improved LLM as a Judge Techniques

Researchers propose BINEVAL, a framework that decomposes LLM evaluation into atomic binary questions for interpretable, multi-dimensional scoring. The method matches or outperforms strong baselines on…

// co-occurs with top 8 entities
// topics top 6 topics