03:41
2026-09-14
zatona.dev
ai-research
The Two MMLU Scores: What a Benchmark Name Does Not Fix
Two MMLU accuracy scores of 0.781 for build 42 and 0.79 for build 44 of the acme-gpt-7b model family are structurally valid but return "incomparable" from the score-delta verifier in the apl-ai-eval cā¦