cd/entity/lm-eval-harnessΒ· homeβ€Ί entitiesβ€Ί lm-eval-harness
grep -l @lm-eval-harness /news/*.json | wc -l β†’ 1

lm-eval-harness

mentions 1 type Organization feed RSS

// recent coverage 1 mentions

15:20
2026-08-20
dev.to
large-language-models

acc vs acc_norm: Why Length Bias Skews LLM Eval Scores

A developer explains how the choice between `acc` and `acc_norm` in lm-eval-harness can skew LLM evaluation results due to length bias. The raw `acc` metric favors shorter answers because it sums toke…

// co-occurs with top 1 entities
// topics top 3 topics