cd/entity/APL AI-EvalĀ· home› entities› APL AI-Eval
grep -l @apl ai-eval /news/*.json | wc -l → 1

APL AI-Eval

mentions 1 type Person feed RSS

// recent coverage 1 mentions

03:41
2026-09-14
zatona.dev
ai-research

The Two MMLU Scores: What a Benchmark Name Does Not Fix

Two MMLU accuracy scores of 0.781 for build 42 and 0.79 for build 44 of the acme-gpt-7b model family are structurally valid but return "incomparable" from the score-delta verifier in the apl-ai-eval c…

// co-occurs with top 5 entities
// topics top 3 topics