cd/entity/AlpacaEval· home entities AlpacaEval
grep -l @alpacaeval /news/*.json | wc -l → 4

AlpacaEval

mentions 4 type Organization feed RSS

// recent coverage 4 mentions

05:17
2026-07-26
kraghavan.ca
large-language-models

LLM-as-a-Judge Field Guide

A field guide to LLM-as-a-judge reveals that five novel proposals for improving the technique were all refuted by prior art published within the last twelve weeks, according to a research synthesis co…

05:00
2026-05-26
alex.smola.org
large-language-models

You don't need all the LLM benchmarks

A new analysis of over 5,400 AI models reveals that benchmark scores for large language models are highly correlated, with just five subjects on the MMLU test predicting the remaining 52 with 91% accu…

// co-occurs with top 8 entities
// topics top 6 topics