cd /news/ai-agents/composite-bench · home topics ai-agents article
[ARTICLE · art-71197] src=benchmarklist.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Composite-Bench

Composite-Bench, a long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments, scores AI agents on pass@1 against certified optimal solutions across five headline slices and a sixth number-partitioning slice with a separate reward ladder. Claude Fable 5 ranks first with a score of 74 on the top imported result.

read1 min views1 publishedJul 24, 2026

Long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments. Its five headline slices score pass@1 against certified optimal solutions, while a sixth number-partitioning slice reports a separate reward ladder for deliberately intractable tasks. Category: Agentic. Imported rows: 8. Top imported result: Claude Fable 5, rank 1, 74.

── more in #ai-agents 4 stories · sorted by recency
── more on @composite-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/composite-bench] indexed:0 read:1min 2026-07-24 ·