cd/entity/Long-Horizon-Terminal-Bench· home entities Long-Horizon-Terminal-Bench
grep -l @long-horizon-terminal-bench /news/*.json | wc -l → 5

Long-Horizon-Terminal-Bench

mentions 5 type Organization feed RSS

// recent coverage 5 mentions

12:01
2026-07-13
dev.to
artificial-intelligence

Top AI Papers on Hugging Face - 2026-07-13

A roundup of top AI papers on Hugging Face highlights four major themes: real-time interactive video generation with Vidu S1, scientific reasoning on molecular structures via SciReasoner, a critique o…

06:38
2026-07-13
machinebrief.com
artificial-intelligence

Are AI Agents Ready for the Long Game?

A new benchmark, Long-Horizon-Terminal-Bench, tests AI agents on 46 long-horizon tasks averaging 231 episodes and 85.3 minutes per run, with top models achieving only a 15.2% pass rate at a 0.95 parti…

// co-occurs with top 8 entities
// topics top 6 topics