04:00
2026-08-18
machinebrief.com
artificial-intelligence
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
A new arXiv paper (2608.14711v1) reveals that AI coding agent benchmarks misapply the pass@k estimator, inflating reported scores by 0.85-0.97 in absolute terms (0.96-0.98 reported vs. 0.00-0.12 correβ¦