cd/entity/MMLU· home› entities› MMLU
grep -l @mmlu /news/*.json | wc -l → 71

MMLU

mentions 71 type Organization page 2/4 feed RSS

// recent coverage 71 mentions

04:00
2026-08-12
arxiv.org
artificial-intelligence

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

Researchers introduced LLM Agents Factory, a retrieval-based framework that constructs domain-specific LLM agents on demand from a base of over 20,000 predetermined agent profiles, achieving accuracy …

15:22
2026-08-04
promptcube3.com
artificial-intelligence

Computer Anthology: The Benchmark That Grows With AI

Computer Anthology, a new benchmark family from the ACL Anthology repository, evaluates AI models on reasoning over newly published NLP papers, with monthly updates to prevent data contamination. The …

04:00
2026-08-04
machinebrief.com
large-language-models

Gaokerena: A Small Persian Medical Language Model Family

Researchers introduced Gaokerena, a family of compact Persian medical language models, with Gaokerena-V improving a translated medical MMLU benchmark from 46.28% to 49.31% and Gaokerena-R achieving 52…

20:44
2026-07-30
lesswrong.com
ai-research

Hint-based CoT faithfulness evals still mostly work on Claude

Redwood Research finds that hint-based chain-of-thought faithfulness evaluations still work on Claude models, contradicting Anthropic system card claims that recent models no longer use hints. The rep…

07:38
2026-07-26
narracomm.com
artificial-intelligence

Best AI Prompts for 2026: What Changed & What Still Works

Three prompting techniques that worked in 2023 now actively cost accuracy in 2026, according to recent research. Expert personas reduced factual accuracy to 68.0% versus 71.6% without on the MMLU know…

04:00
2026-07-21
arxiv.org
artificial-intelligence

Diagnosing Correctness Probes under Self-Judgement Confounding

A new arXiv preprint (2607.16799v1) finds that hidden-state readouts from language models primarily encode self-judgement (SJ) rather than objective correctness (OC), with the SJ-associated direction …

00:00
2026-07-19
zackproser.com
artificial-intelligence

The Benchmark

A benchmark score is a manufactured number that passes through a chain of choices—sampling, prompting, scoring, and aggregation—each of which can change the final result while model weights stay fixed…

← prev page 2 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics