Composite-Bench
Composite-Bench, a long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments, scores AI agents on pass@1 against certified optimal solutions across five he…
Composite-Bench, a long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments, scores AI agents on pass@1 against certified optimal solutions across five he…
Founder Bench (eico.so "acoco" live arena) pits five frontier models against each other by giving each four identical businesses to operate autonomously in the real world, with no human intervention. …
BenchmarkList, a new platform tracking over 2,400 AI benchmarks, models, and capabilities, launched to centralize scattered evaluation data from research papers, GitHub repos, model cards, and tweets.…
WANDR (Wide ANd Deep Research) evaluates research agents on 500 realistic data-collection tasks requiring broad entity discovery and evidence-backed claims, using a reference-free grader that re-fetch…
DuelLab GameBench 2 evaluates AI-generated game-playing programs by compiling and running them head-to-head across an evolving set of named public board and strategy games, with Claude Fable 5 ranking…
FrontierFinance, a new benchmark from an unnamed source, evaluates AI research systems on 220 open-ended financial research queries using 11,543 expert-authored rubric criteria across six workflows. T…
A new benchmark, Long-Horizon Terminal-Bench, evaluates AI agents on 46 terminal tasks across nine categories, including software engineering and scientific computing, using a 90-minute budget and hid…
Researchers introduced Vera, an automated safety testing framework for LLM agents, and released Vera-Bench, a benchmark of 1600 executable safety cases. Testing on four production agent frameworks rev…
Researchers introduced MindEdit-Bench, a benchmark of six spatial reasoning tasks for vision-language models, built from smartphone triplets of 120 private indoor scenes. Across 15 VLMs on 1,003 human…
Researchers released GeneBench-Pro, a 129-problem benchmark testing AI agents on multi-stage genomics and biomedicine analyses. GPT-5.6 Sol Pro (Extended) achieved the top score of 31.5%, highlighting…
MemDelta, a controlled evaluation protocol for agent memory systems, reveals that reported performance gains often stem from changes in embedding models or language models rather than memory architect…
Researchers introduced Cortex, a framework that organizes web-scale corpora into an Ontological Corpus Graph (OCG) for structured knowledge organization, addressing the lack of systematic organization…
Researchers introduced the Human Creativity Benchmark (HCB), a new evaluation framework that separates professional agreement on technical dimensions from legitimate disagreement on taste-driven dimen…
Researchers introduced MuseBench, a benchmark with 4,016 questions evaluating multimodal large language models on intent-level understanding of audiovisual arts across cinema, visual arts, stage perfo…
Researchers at an undisclosed institution introduced clinical reasoning graphs to evaluate LLM diagnostic reasoning across 750 traces from five models on 50 NEJM cases. They found that while LLMs achi…
Researchers at an undisclosed institution developed MemLeak, a benchmark revealing that multimodal AI agents fail to fully delete user facts because retained images allow visual information recovery. …
Researchers introduced CLQT, a closed-loop evaluation framework for LLM-based trading agents that diagnoses reasoning and strategy consistency rather than ranking by returns. The system uses a five-st…
Researchers introduced NL-PDDL-Bench, a benchmark for converting natural language into PDDL planning specifications, and a planner-in-the-loop framework that improves executability and safety. Their m…