cd/entity/HumanEval· home› entities› HumanEval
grep -l @humaneval /news/*.json | wc -l → 57

HumanEval

mentions 57 type Organization page 3/3 feed RSS

// recent coverage 57 mentions

05:00
2026-06-29
dev.to
artificial-intelligence

AI/ML Research Digest — Jun 27, 2026

Recent AI research introduces RL-driven agentic optimization using dense token-level supervision and progress advantage signals to stabilize training. PhysiFormer injects 3D geometric reasoning into d…

23:17
2026-06-18
blog.r-lopes.com
ai-safety

The Line Vibe Coding Can't Cross

Vibe coding—prompting an AI agent and shipping unread output—introduces a measurable defect tax that makes it unsuitable for mission-critical systems, with 45% of AI-generated code containing security…

16:07
2026-06-17
danlevy.net
large-language-models

LLM benchmarks are answering someone else's question

LLM benchmarks like MMLU and HumanEval are irrelevant for most businesses building AI products, as they measure generic performance rather than specific system tasks. Teams should instead build custom…

15:03
2026-06-17
dev.to
large-language-models

Claude 3.5 Sonnet Isn't Just an Upgrade. It's a New Baseline.

Anthropic released Claude 3.5 Sonnet, a new AI model that outperforms the previous top-tier Claude 3 Opus in intelligence, speed, and cost. The model achieves a 64% solve rate on internal agentic codi…

04:59
2026-06-17
dev.to
large-language-models

Kog hits 3K t/s on MI300X, no kernel switches — test it now

Kog AI achieved over 3,000 output tokens per second per request for an FP16 2B model on a single 8× MI300X node using a monokernel that eliminates per-token kernel launches. The technique collapses th…

09:35
2026-06-14
dev.to
large-language-models

Running Chinese LLMs at Scale: A Cloud Architect's Notes

A cloud architect evaluated four Chinese LLM families—DeepSeek, Qwen, Kimi, and GLM—in a multi-region production pipeline serving thousands of requests per second via Global API's unified endpoint. De…

20:09
2026-06-04
mendral.com
ai-agents

How we know if our agent is right

Mendral, an AI DevOps agent developer, cannot provide a single accuracy metric for its CI failure diagnosis agent despite processing 36,564 investigations across 5.7 million CI jobs and 14.4 billion l…

17:11
2026-05-30
github.com
ai-agents

The future will be millions agents running task everyday?

A new benchmark comparing agent runtime performance across C++, Python, TypeScript, and Rust found that C++ achieved a peak memory footprint of approximately 93 MiB while running 100 concurrent coding…

15:19
2026-05-20
dev.to
large-language-models

What did gemma see? - Thinking in comments...

The Gemma 4 26B model was the first local AI to achieve a perfect score on the HumanEval benchmark, including solving the notoriously difficult problem 145. This problem requires sorting integers by t…

← prev page 3 / 3
// co-occurs with top 8 entities
// topics top 6 topics