cd/entity/MMLU· home› entities› MMLU
grep -l @mmlu /news/*.json | wc -l → 71

MMLU

mentions 71 type Organization page 1/4 feed RSS

// recent coverage 71 mentions

20:00
2026-09-19
jdhornsby.com
artificial-intelligence

Fifty cents of Jev

TypeSafe released a new "System One" model called Jev that answers typed questions with a chosen answer, per-option probabilities, and a confidence score rather than generating text, priced at $0.042 …

07:00
2026-09-18
hamel.dev
ai-research

AI Evals: Everything You Need to Know

Hamel Husain and Shreya Shankar published an AI Evals FAQ distilling the most common questions from teaching 700+ engineers and product managers about AI evaluation. The guide distinguishes model benc…

03:41
2026-09-14
zatona.dev
ai-research

The Two MMLU Scores: What a Benchmark Name Does Not Fix

Two MMLU accuracy scores of 0.781 for build 42 and 0.79 for build 44 of the acme-gpt-7b model family are structurally valid but return "incomparable" from the score-delta verifier in the apl-ai-eval c…

04:09
2026-09-02
promptcube3.com
large-language-models

Why are LLM benchmarks looking so completely unhinged lately

LLM benchmark scores are becoming unreliable because models are trained on the same tests they are evaluated on, leading to inflated results that fail to reflect real-world performance, according to a…

20:54
2026-08-31
promptcube3.com
large-language-models

Why Qwen3.

A hands-on benchmark of Qwen3.8 27B on an RTX 4090 found that 4-bit quantization (NF4 and AWQ INT4) preserves near-baseline quality, with MMLU scores of 59.8% and 59.5% versus 61.2% for FP16, while 1-…

18:01
2026-08-21
promptcube3.com
artificial-intelligence

COPA treats prompt injection as lifelong learning not one-time

A new method called COPA reduces attack success rates by 6.3× versus the best static baseline and 4.4× on average across lifelong attack streams, while retaining 92% defense on month-old attacks compa…

15:14
2026-08-20
promptcube3.com
artificial-intelligence

Rethinking LLM scaling after Jie Tang's latest breakdown

Jie Tang's team at Z.ai argues that LLM scaling should optimize for fixed inference budgets, favoring smaller, deeper architectures trained longer over wider models. Their ablation shows a 7B model tr…

page 1 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics