cd/entity/Terminal Bench· home entities Terminal Bench
grep -l @terminal bench /news/*.json | wc -l → 10

Terminal Bench

mentions 10 type Person feed RSS

// recent coverage 10 mentions

00:00
2026-09-02
mindstudio.ai
artificial-intelligence

Claude Fable 5.1: What's New in Anthropic's Latest Model

Anthropic released Claude Fable 5.1, a refinement of its Fable 5 model that improves agentic performance and readability while cutting cached prompt read costs by about 75%, resulting in roughly 25% c…

18:10
2026-08-18
cline.ghost.io
ai-research

Open-sourcing evals for open-weight agents

Cline, an AI coding assistant, is open-sourcing its evaluation framework for open-weight coding agents, revealing that its requests run 20–30% heavier on tokens than the most efficient harnesses. The …

00:00
2026-08-04
mindstudio.ai
artificial-intelligence

Qwen 3.8 Max Explained: Alibaba's 2.4 Trillion Parameter Model

Alibaba released Qwen 3.8 Max, a 2.4 trillion parameter open-weight AI model with 95 billion active parameters, set to become the largest open-weight model once weights are open-sourced about a week a…

15:29
2026-07-14
forum.effectivealtruism.org
ai-safety

Good Benchmarks

METR contributor Ivan Bercovich argues that most AI benchmarks are flawed and that building good ones requires nuanced understanding, drawing on 18 months of experience with Terminal Bench. Good tasks…

13:12
2026-06-19
blog.kilo.ai
ai-agents

Terminal Bench Scores Are Now in Your Editor

Terminal Bench completion scores and per-attempt costs now appear directly in the Kilo CLI and VS Code extension model details panel, providing real benchmark data where developers choose models. The …

// co-occurs with top 8 entities
// topics top 6 topics