cd/entity/MMLU· home› entities› MMLU
grep -l @mmlu /news/*.json | wc -l → 71

MMLU

mentions 71 type Organization page 3/4 feed RSS

// recent coverage 71 mentions

17:08
2026-07-10
machinebrief.com
artificial-intelligence

Breaking Down Long-Context Transformer Bottlenecks

Researchers have developed a new approach to overcome the quadratic cost of causal self-attention in long-context transformers, using state update design and structural interventions like sink tokens …

06:27
2026-07-09
letsdatascience.com
artificial-intelligence

AI Benchmark Scores Overstate Model Performance

A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…

00:17
2026-07-07
supercomputing-system-ai-lab.github.io
machine-learning

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models

Researchers introduce PuzzleMoE, a method for compressing large Mixture-of-Experts models via fine-grained element-wise merging and bit-packing, achieving up to 16.7% higher accuracy on MMLU at 50% co…

12:00
2026-07-02
kdnuggets.com
artificial-intelligence

Humanity’s Last Exam is a Distraction

The Center for AI Safety, with world experts, created Humanity's Last Exam (HLE), a benchmark of over 2,500 expert-level questions across disciplines to test AI reasoning. Even top models like GPT, Ge…

05:00
2026-07-01
dev.to
machine-learning

RL-driven data mixing boosts evaluation scores

A reinforcement learning-driven data scheduler, AC-ODM, boosts MMLU performance by 27.5% relative and HumanEval pass@1 by 2.23× on a Pythia-1B model with only a 0.4% per-step wall-clock increase and 2…

16:01
2026-06-30
pub.towardsai.net
ai-agents

AI Agent Evaluation: How to Know If Your Agent Actually Works

A developer recounts pushing an agent into production that failed after a CRM dropdown change, highlighting the inadequacy of model-level evaluation for agent systems. The article argues that agents m…

00:00
2026-06-30
huggingface.co
ai-research

Featuring Every Eval Ever Results on Hugging Face Model Pages

Hugging Face and the EvalEval Coalition launched an integration that allows contributors to submit standardized evaluation results (EEE schema) to Hugging Face Community Evals, consolidating scattered…

00:00
2026-06-30
evalevalai.com
artificial-intelligence

When AI Benchmarks Stop Measuring Progress

Nearly half of 60 popular AI benchmarks studied in a new ICML paper show high levels of saturation, meaning they can no longer reliably distinguish between leading models. The researchers from the pap…

05:00
2026-06-29
dev.to
artificial-intelligence

AI/ML Research Digest — Jun 27, 2026

Recent AI research introduces RL-driven agentic optimization using dense token-level supervision and progress advantage signals to stabilize training. PhysiFormer injects 3D geometric reasoning into d…

00:04
2026-06-27
devclubhouse.com
artificial-intelligence

The Open-Weights Gap Depends on What You Measure

A viral chart predicting open-weights AI models will catch closed frontier models by December 2026 is misleading, as analysis across 18 benchmarks shows the gap varies by task and is not uniformly shr…

12:04
2026-06-25
discuss.huggingface.co
machine-learning

What's your method for benchmarking?

A practical guide for benchmarking fine-tuned models recommends starting with a held-out test set matching the actual task rather than relying solely on public benchmarks. The workflow includes defini…

00:00
2026-06-24
jasonrobert.dev
large-language-models

News Summary for June 24, 2026

Microsoft Principal Developer Advocate Waldek Mastykarz published a blog post arguing that testing large language models in empty chat sessions is methodologically flawed, as models have no preference…

14:53
2026-06-22
devblogs.microsoft.com
large-language-models

Models don’t have preferences, they have context

A new analysis argues that claims about large language models having preferences, such as 'Claude prefers React,' are misleading because models lack preferences and instead reflect training data and c…

04:24
2026-06-21
thewatershed.markpesce.com
ai-safety

Why Evals are Hard

AI evaluations are failing as models approach general intelligence, with benchmarks saturating through contamination and Goodhart effects while the scope of evaluation expands from minutes to months. …

16:18
2026-06-18
lesswrong.com
ai-safety

Your Model Organisms Might Be Fried

Arcadia Alignment's research reveals that current AI model organisms used to study alignment pathologies suffer from degraded coherence, instruction-following, and reasoning, making them poor proxies …

← prev page 3 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics