cd/entity/MMLU· home entities MMLU
grep -l @mmlu /news/*.json | wc -l → 56

MMLU

mentions 56 type Organization page 2/3 feed RSS

// recent coverage 56 mentions

04:00
2026-07-21
arxiv.org
artificial-intelligence

Diagnosing Correctness Probes under Self-Judgement Confounding

A new arXiv preprint (2607.16799v1) finds that hidden-state readouts from language models primarily encode self-judgement (SJ) rather than objective correctness (OC), with the SJ-associated direction …

00:00
2026-07-19
zackproser.com
artificial-intelligence

The Benchmark

A benchmark score is a manufactured number that passes through a chain of choices—sampling, prompting, scoring, and aggregation—each of which can change the final result while model weights stay fixed…

17:08
2026-07-10
machinebrief.com
artificial-intelligence

Breaking Down Long-Context Transformer Bottlenecks

Researchers have developed a new approach to overcome the quadratic cost of causal self-attention in long-context transformers, using state update design and structural interventions like sink tokens …

06:27
2026-07-09
letsdatascience.com
artificial-intelligence

AI Benchmark Scores Overstate Model Performance

A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…

00:17
2026-07-07
supercomputing-system-ai-lab.github.io
machine-learning

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models

Researchers introduce PuzzleMoE, a method for compressing large Mixture-of-Experts models via fine-grained element-wise merging and bit-packing, achieving up to 16.7% higher accuracy on MMLU at 50% co…

12:00
2026-07-02
kdnuggets.com
artificial-intelligence

Humanity’s Last Exam is a Distraction

The Center for AI Safety, with world experts, created Humanity's Last Exam (HLE), a benchmark of over 2,500 expert-level questions across disciplines to test AI reasoning. Even top models like GPT, Ge…

05:00
2026-07-01
dev.to
machine-learning

RL-driven data mixing boosts evaluation scores

A reinforcement learning-driven data scheduler, AC-ODM, boosts MMLU performance by 27.5% relative and HumanEval pass@1 by 2.23× on a Pythia-1B model with only a 0.4% per-step wall-clock increase and 2…

16:01
2026-06-30
pub.towardsai.net
ai-agents

AI Agent Evaluation: How to Know If Your Agent Actually Works

A developer recounts pushing an agent into production that failed after a CRM dropdown change, highlighting the inadequacy of model-level evaluation for agent systems. The article argues that agents m…

00:00
2026-06-30
huggingface.co
ai-research

Featuring Every Eval Ever Results on Hugging Face Model Pages

Hugging Face and the EvalEval Coalition launched an integration that allows contributors to submit standardized evaluation results (EEE schema) to Hugging Face Community Evals, consolidating scattered…

00:00
2026-06-30
evalevalai.com
artificial-intelligence

When AI Benchmarks Stop Measuring Progress

Nearly half of 60 popular AI benchmarks studied in a new ICML paper show high levels of saturation, meaning they can no longer reliably distinguish between leading models. The researchers from the pap…

05:00
2026-06-29
dev.to
artificial-intelligence

AI/ML Research Digest — Jun 27, 2026

Recent AI research introduces RL-driven agentic optimization using dense token-level supervision and progress advantage signals to stabilize training. PhysiFormer injects 3D geometric reasoning into d…

00:04
2026-06-27
devclubhouse.com
artificial-intelligence

The Open-Weights Gap Depends on What You Measure

A viral chart predicting open-weights AI models will catch closed frontier models by December 2026 is misleading, as analysis across 18 benchmarks shows the gap varies by task and is not uniformly shr…

12:04
2026-06-25
discuss.huggingface.co
machine-learning

What's your method for benchmarking?

A practical guide for benchmarking fine-tuned models recommends starting with a held-out test set matching the actual task rather than relying solely on public benchmarks. The workflow includes defini…

00:00
2026-06-24
jasonrobert.dev
large-language-models

News Summary for June 24, 2026

Microsoft Principal Developer Advocate Waldek Mastykarz published a blog post arguing that testing large language models in empty chat sessions is methodologically flawed, as models have no preference…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics