cd/sources/benchmarklist-auto-discovered· home sources Benchmarklist (auto-discovered)
cat /sources/benchmarklist-auto-discovered.feed | wc -l → 18

Benchmarklist (auto-discovered)

articles 18 domain benchmarklist.com → feed RSS
00:00
2026-07-24
benchmarklist.com
ai-agents

Composite-Bench

Composite-Bench, a long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments, scores AI agents on pass@1 against certified optimal solutions across five he…

00:00
2026-07-16
benchmarklist.com
ai-agents

Founder Bench

Founder Bench (eico.so "acoco" live arena) pits five frontier models against each other by giving each four identical businesses to operate autonomously in the real world, with no human intervention. …

00:00
2026-07-14
benchmarklist.com
ai-research

WANDR

WANDR (Wide ANd Deep Research) evaluates research agents on 500 realistic data-collection tasks requiring broad entity discovery and evidence-backed claims, using a reference-free grader that re-fetch…

00:00
2026-07-13
benchmarklist.com
artificial-intelligence

DuelLab GameBench 2

DuelLab GameBench 2 evaluates AI-generated game-playing programs by compiling and running them head-to-head across an evolving set of named public board and strategy games, with Claude Fable 5 ranking…

00:00
2026-07-13
benchmarklist.com
artificial-intelligence

FrontierFinance

FrontierFinance, a new benchmark from an unnamed source, evaluates AI research systems on 220 open-ended financial research queries using 11,543 expert-authored rubric criteria across six workflows. T…

00:00
2026-07-09
benchmarklist.com
ai-agents

Long-Horizon Terminal-Bench

A new benchmark, Long-Horizon Terminal-Bench, evaluates AI agents on 46 terminal tasks across nine categories, including software engineering and scientific computing, using a 90-minute budget and hid…

00:00
2026-07-02
benchmarklist.com
ai-safety

Vera-Bench

Researchers introduced Vera, an automated safety testing framework for LLM agents, and released Vera-Bench, a benchmark of 1600 executable safety cases. Testing on four production agent frameworks rev…

00:00
2026-07-01
benchmarklist.com
computer-vision

MindEdit-Bench

Researchers introduced MindEdit-Bench, a benchmark of six spatial reasoning tasks for vision-language models, built from smartphone triplets of 120 private indoor scenes. Across 15 VLMs on 1,003 human…

00:00
2026-06-30
benchmarklist.com
artificial-intelligence

GeneBench-Pro

Researchers released GeneBench-Pro, a 129-problem benchmark testing AI agents on multi-stage genomics and biomedicine analyses. GPT-5.6 Sol Pro (Extended) achieved the top score of 31.5%, highlighting…

00:00
2026-06-29
benchmarklist.com
artificial-intelligence

MemDelta

MemDelta, a controlled evaluation protocol for agent memory systems, reveals that reported performance gains often stem from changes in embedding models or language models rather than memory architect…

00:00
2026-06-29
benchmarklist.com
artificial-intelligence

Human Creativity Benchmark

Researchers introduced the Human Creativity Benchmark (HCB), a new evaluation framework that separates professional agreement on technical dimensions from legitimate disagreement on taste-driven dimen…

00:00
2026-06-29
benchmarklist.com
artificial-intelligence

CLQT

Researchers introduced CLQT, a closed-loop evaluation framework for LLM-based trading agents that diagnoses reasoning and strategy consistency rather than ranking by returns. The system uses a five-st…

00:00
2026-06-29
benchmarklist.com
artificial-intelligence

NL-PDDL-Bench

Researchers introduced NL-PDDL-Bench, a benchmark for converting natural language into PDDL planning specifications, and a planner-in-the-loop framework that improves executability and safety. Their m…