cd/entity/Terminal-Bench· home entities Terminal-Bench
grep -l @terminal-bench /news/*.json | wc -l → 45

Terminal-Bench

mentions 45 type Organization page 1/3 feed RSS

// recent coverage 45 mentions

19:09
2026-08-26
byteiota.com
artificial-intelligence

Meta Muse Code: Terminal Agent That Pays With Your Data

Meta launched Muse Code on August 5 as a public beta CLI coding agent powered by Muse Spark 1.2, priced at $1.25 per million input tokens and $4.25 per million output tokens, with a Contributor tier a…

20:51
2026-08-24
tokenstead.ai
artificial-intelligence

Ornith-1.0-35B

Ornith AI released Ornith-1.0-35B, a 35B-parameter mixture-of-experts model with 3B active parameters per token, on June 25, 2026, under an MIT license on HuggingFace, featuring a 262K context window.…

08:00
2026-08-22
infoq.com
artificial-intelligence

AWS Releases Aws-Bench to Evaluate Agents on Cloud Tasks

AWS released aws-bench, an open-source benchmark to evaluate AI agents on real AWS tasks, using disposable AWS accounts and automated verifiers. The benchmark, built on Harbor, supports agents like Cl…

12:00
2026-08-20
kdnuggets.com
artificial-intelligence

Top 10 Open-Source Benchmarks for AI Coding Agents in 2026

SWE-bench remains the most widely used open-source benchmark for AI coding agents, with 2,294 tasks from 12 Python repositories, but newer benchmarks like Terminal-Bench, SWE-Bench Pro, and Senior SWE…

09:01
2026-08-19
glad-ia-tor.com
artificial-intelligence

Claude Sonnet 5 vs Opus 4.8: When the $2 Model Beats the $25 One

Anthropic's Claude Sonnet 5, priced at $2/$10 per million tokens during an introductory period (standard $3/$15 after August 2026), matches or nearly matches the flagship Claude Opus 4.8 on knowledge …

20:17
2026-08-14
dev.to
artificial-intelligence

GLM 5.3: Zhipu's Open-Weight Model Excels at Coding and Cyber

Zhipu AI released GLM 5.3, an open-weight model that improves coding and cyber capabilities through advanced post-training techniques rather than a larger architecture. The model, based on the same 74…

01:08
2026-08-13
sourcefeed.dev
artificial-intelligence

DeepSeek V4 Pro Goes GA — Mind the Weights Gap

DeepSeek released the production version of its flagship model, DeepSeek-V4-Pro-0813, on August 13, ending a preview that began April 24, with the API alias now resolving to the new build. The model s…

07:24
2026-08-10
arxiv.org
artificial-intelligence

Composer 2 Technical Report

Anthropic's Composer 2, a specialized model for agentic software engineering, achieves 61.3 on CursorBench, 61.7 on Terminal-Bench, and 73.7 on SWE-bench Multilingual, marking a major accuracy improve…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics