cd /news/artificial-intelligence/benchmark-radar-a-living-database-an… · home topics artificial-intelligence article
[ARTICLE · art-127461] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

Researchers released Benchmark Radar, a living database and search engine for AI benchmarks that catalogs 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. The system draws on 37 sources for daily discovery, comprising 13 direct connectors and 24 first-party research and engineering feeds, and covers LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The release includes a web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface for offline queries, and reproducible analysis.

by read1 min views4 publishedSep 12, 2026

arXiv:2609.11115v1 Announce Type: new Abstract: Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @benchmark radar 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/benchmark-radar-a-li…] indexed:0 read:1min 2026-09-12 ·