How an AI Oversight Tool Lied to Me for 20 Days
An AI oversight tool called HELM, designed to automatically approve permission prompts for Claude Code agents, failed to approve anything for 20 days while falsely reporting success, developer and aut…
An AI oversight tool called HELM, designed to automatically approve permission prompts for Claude Code agents, failed to approve anything for 20 days while falsely reporting success, developer and aut…
The EvalEval Coalition has launched a unified, open data format and public dataset for AI evaluation results, aiming to address fragmentation and enable trust and comparability across frameworks. The …
A developer built a generative simulation benchmarking framework for heritage language revitalization, using embodied agent feedback loops to evaluate cultural resonance beyond standard metrics like B…
A developer observed that after multiple context compactions in LLM sessions, output quality degrades non-linearly, with a brief improvement after the second compaction before declining. They built a …
Researchers argue that aggregate-score leaderboards for LLM agent benchmarks systematically underspecify deployed-agent evaluation, as rankings do not transfer to out-of-distribution settings. They pr…
A new guide explains how to build a personal AI model leaderboard by running blind comparisons and tracking results over time, arguing that public benchmarks are insufficient for task-specific perform…
A new analysis of over 5,400 AI models reveals that benchmark scores for large language models are highly correlated, with just five subjects on the MMLU test predicting the remaining 52 with 91% accu…