cd/entity/HellaSwag· home entities HellaSwag
grep -l @hellaswag /news/*.json | wc -l → 10

HellaSwag

mentions 10 type Organization feed RSS

// recent coverage 10 mentions

15:20
2026-08-20
dev.to
large-language-models

acc vs acc_norm: Why Length Bias Skews LLM Eval Scores

A developer explains how the choice between `acc` and `acc_norm` in lm-eval-harness can skew LLM evaluation results due to length bias. The raw `acc` metric favors shorter answers because it sums toke…

06:27
2026-07-09
letsdatascience.com
artificial-intelligence

AI Benchmark Scores Overstate Model Performance

A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…

00:00
2026-06-09
andlukyane.com
machine-learning

Book Review: 50 ML Projects to Understand LLMs

Mike X Cohen's new book "50 ML Projects to Understand LLMs" uses GPT-2 as a scientific specimen, teaching readers to investigate the model through 50 hands-on projects focused on code, statistics, and…

// co-occurs with top 8 entities
// topics top 6 topics