cd/entity/HumanEval· home entities HumanEval
grep -l @humaneval /news/*.json | wc -l → 42

HumanEval

mentions 42 type Organization page 1/3 feed RSS

// recent coverage 42 mentions

01:00
2026-08-22
smarterarticles.co.uk
artificial-intelligence

95% Solved: Why AI Code Still Ignores Your Instructions

A team led by Ming Zhong at the University of Illinois Urbana-Champaign and Google DeepMind has introduced SWE-IF, a framework that aligns code evaluation with human preference, revealing that instruc…

15:14
2026-08-20
promptcube3.com
artificial-intelligence

Rethinking LLM scaling after Jie Tang's latest breakdown

Jie Tang's team at Z.ai argues that LLM scaling should optimize for fixed inference budgets, favoring smaller, deeper architectures trained longer over wider models. Their ablation shows a 7B model tr…

04:00
2026-08-20
machinebrief.com
artificial-intelligence

What is Missing from AI Post-Training AI: An Empirical Analysis

A new empirical analysis from arXiv (2608.19072v1) finds that large language model (LLM) agents post-training an LLM lock in their training strategy at the very beginning and spend the remaining budge…

00:00
2026-08-09
bharad.dev
artificial-intelligence

Measuring Agent Reliability: pass@k, pass^k, and LLM Judges

An agent that passes 75 out of 100 test runs has a 75% pass@1 rate, but reliability depends on the product: pass@k (at least one success in k tries) yields 98.4% for k=3, while pass^k (all k succeed) …

11:13
2026-08-06
machinelearningmastery.com
artificial-intelligence

Designing AI Agents That Can Self-Correct

A new tutorial from Anthropic demonstrates that AI agents can reliably self-correct only when grounded in external verification, citing a 2024 paper showing large language models cannot self-correct r…

18:09
2026-07-23
promptcube3.com
artificial-intelligence

DeepSeek V3 vs Claude for coding

DeepSeek V3 generally outperforms Claude 3.5 Sonnet on raw coding benchmarks like HumanEval and MBPP, especially in C++, Rust, and mathematical implementations, while Claude 3.5 Sonnet is considered s…

01:54
2026-07-16
dev.to
large-language-models

LLM Evals For Developer Tools: Useful, Correct, Safe

A developer argues that LLM evaluations for developer tools must measure three distinct axes: usefulness, correctness, and safety. Unlike chatbot evals, dev-tool evals rely on ground-truth signals lik…

06:27
2026-07-09
letsdatascience.com
artificial-intelligence

AI Benchmark Scores Overstate Model Performance

A PlainEnglish article warns that AI benchmark scores such as MMLU, HumanEval, and HellaSwag can overstate production readiness when leaderboard numbers are treated as proof of model quality. The comm…

01:11
2026-07-08
byteiota.com
large-language-models

NVIDIA Nemotron TwoTower: Run LLMs 2.42x Faster Now

NVIDIA open-sourced Nemotron-Labs-TwoTower, a diffusion language model that generates text 2.42x faster than its autoregressive counterpart without retraining original weights. The model achieves 98.7…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics