cd/entity/MMLU· home entities MMLU
grep -l @mmlu /news/*.json | wc -l → 56

MMLU

mentions 56 type Organization page 3/3 feed RSS

// recent coverage 56 mentions

14:53
2026-06-22
devblogs.microsoft.com
large-language-models

Models don’t have preferences, they have context

A new analysis argues that claims about large language models having preferences, such as 'Claude prefers React,' are misleading because models lack preferences and instead reflect training data and c…

04:24
2026-06-21
thewatershed.markpesce.com
ai-safety

Why Evals are Hard

AI evaluations are failing as models approach general intelligence, with benchmarks saturating through contamination and Goodhart effects while the scope of evaluation expands from minutes to months. …

16:18
2026-06-18
lesswrong.com
ai-safety

Your Model Organisms Might Be Fried

Arcadia Alignment's research reveals that current AI model organisms used to study alignment pathologies suffer from degraded coherence, instruction-following, and reasoning, making them poor proxies …

16:07
2026-06-17
danlevy.net
large-language-models

LLM benchmarks are answering someone else's question

LLM benchmarks like MMLU and HumanEval are irrelevant for most businesses building AI products, as they measure generic performance rather than specific system tasks. Teams should instead build custom…

04:00
2026-05-29
arxiv.org
large-language-models

Mind Your Tone: Does Tone Alter LLM Performance?

A new study from arXiv (2605.29027) found that tonal variations in prompts cause systematic but model-dependent accuracy shifts in large language models (LLMs) on objective multiple-choice questions. …

05:00
2026-05-26
alex.smola.org
large-language-models

You don't need all the LLM benchmarks

A new analysis of over 5,400 AI models reveals that benchmark scores for large language models are highly correlated, with just five subjects on the MMLU test predicting the remaining 52 with 91% accu…

← prev page 3 / 3
// co-occurs with top 8 entities
// topics top 6 topics