cd/entity/HumanEval· home› entities› HumanEval
grep -l @humaneval /news/*.json | wc -l → 57

HumanEval

mentions 57 type Organization page 2/3 feed RSS

// recent coverage 57 mentions

00:00
2026-08-09
bharad.dev
artificial-intelligence

Measuring Agent Reliability: pass@k, pass^k, and LLM Judges

An agent that passes 75 out of 100 test runs has a 75% pass@1 rate, but reliability depends on the product: pass@k (at least one success in k tries) yields 98.4% for k=3, while pass^k (all k succeed) …

11:13
2026-08-06
machinelearningmastery.com
artificial-intelligence

Designing AI Agents That Can Self-Correct

A new tutorial from Anthropic demonstrates that AI agents can reliably self-correct only when grounded in external verification, citing a 2024 paper showing large language models cannot self-correct r…

18:09
2026-07-23
promptcube3.com
artificial-intelligence

DeepSeek V3 vs Claude for coding

DeepSeek V3 generally outperforms Claude 3.5 Sonnet on raw coding benchmarks like HumanEval and MBPP, especially in C++, Rust, and mathematical implementations, while Claude 3.5 Sonnet is considered s…

01:54
2026-07-16
dev.to
large-language-models

LLM Evals For Developer Tools: Useful, Correct, Safe

A developer argues that LLM evaluations for developer tools must measure three distinct axes: usefulness, correctness, and safety. Unlike chatbot evals, dev-tool evals rely on ground-truth signals lik…

06:27
2026-07-09
letsdatascience.com
artificial-intelligence

AI Benchmark Scores Overstate Model Performance

A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…

01:11
2026-07-08
byteiota.com
large-language-models

NVIDIA Nemotron TwoTower: Run LLMs 2.42x Faster Now

NVIDIA open-sourced Nemotron-Labs-TwoTower, a diffusion language model that generates text 2.42x faster than its autoregressive counterpart without retraining original weights. The model achieves 98.7…

13:06
2026-07-02
sourcefeed.dev
artificial-intelligence

Beyond Bug Fixing: The Rise of Senior-Level AI Coding Benchmarks

New AI coding benchmarks, including Snorkel AI's Senior SWE-Bench and Scale AI's SWE-Bench Pro, are replacing older evaluations like HumanEval and original SWE-bench to test senior-level engineering s…

05:00
2026-07-01
dev.to
machine-learning

RL-driven data mixing boosts evaluation scores

A reinforcement learning-driven data scheduler, AC-ODM, boosts MMLU performance by 27.5% relative and HumanEval pass@1 by 2.23× on a Pythia-1B model with only a 0.4% per-step wall-clock increase and 2…

15:17
2026-06-30
byteiota.com
large-language-models

Gemini 2.5 Pro Deep Think: What the Benchmarks Mean

Google's Gemini 2.5 Pro with Deep Think reasoning mode topped coding and reasoning benchmarks this week, scoring 82.4% on GPQA Diamond and 94.1% on HumanEval+, but the mode multiplies token costs by r…

00:00
2026-06-30
evalevalai.com
artificial-intelligence

When AI Benchmarks Stop Measuring Progress

Nearly half of 60 popular AI benchmarks studied in a new ICML paper show high levels of saturation, meaning they can no longer reliably distinguish between leading models. The researchers from the pap…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics