cd/entity/CEO-Bench· home entities CEO-Bench
grep -l @ceo-bench /news/*.json | wc -l → 3

CEO-Bench

mentions 3 type Organization feed RSS

// recent coverage 3 mentions

04:00
2026-06-18
arxiv.org
large-language-models

CEO-Bench: Can Agents Play the Long Game?

Researchers introduced CEO-Bench, a benchmark that evaluates language model agents on long-horizon tasks by simulating operating a startup for 500 days. The strongest agents, including Claude Opus 4.8…

// co-occurs with top 8 entities
// topics top 4 topics