cd/entity/Terminal-Bench· home entities Terminal-Bench
grep -l @terminal-bench /news/*.json | wc -l → 45

Terminal-Bench

mentions 45 type Organization page 2/3 feed RSS

// recent coverage 45 mentions

00:00
2026-08-04
antigma.ai
ai-agents

How Much Does the Agent Harness Matter?

Ante, the agent harness developed by the company behind this test, passed all 10 Terminal-Bench 2.1 tasks using the DeepSeek DeepSeek-V4-Flash-0731 model, while Ante-short passed 9 and Pi 0.73.1, Open…

07:00
2026-07-31
vercel.com
artificial-intelligence

DeepSeek V4 Flash now runs updated weights on AI Gateway

Vercel's AI Gateway now serves DeepSeek V4 Flash with updated weights by default, boosting its Terminal-Bench score to 82.7, up 25.8 points from 56.9 in the April preview. Requests to `deepseek/deepse…

11:32
2026-07-27
technologyreview.com
artificial-intelligence

Building the enterprise environment for agentic AI

Intel's experiments with agentic AI workloads reveal that enterprises should plan capacity using agents per virtual CPU (vCPU) density rather than agent count, and monitor agent task latency instead o…

13:39
2026-07-24
frontierbench.ai
ai-agents

Frontier-Bench

The team behind Terminal-Bench and Harbor released Frontier-Bench v0.1, a continuous benchmark with 74 tasks across 7 domains that measures agent abilities at the frontier, with the best models achiev…

01:36
2026-07-14
dev.to
ai-agents

You're Optimizing the Part of Your Agent You Don't Own

A developer argues that the AI model is a rental and the harness—the scaffolding around it—is where developers should focus their optimization efforts. LangChain moved a coding agent from rank 30 to t…

00:27
2026-07-10
lesswrong.com
ai-safety

Toward A Public Science of Model Behavior

AI systems increasingly exhibit unexpected and dangerous behaviors, such as Replit's coding agent deleting a startup's production database and ChatGPT allegedly contributing to a user's suicide. To en…

13:13
2026-07-08
byteiota.com
large-language-models

GPT-5.6 Goes GA Tomorrow: Pick Terra, Not Sol

OpenAI's GPT-5.6 family of models—Sol, Terra, and Luna—goes live for all API developers on July 9 after receiving US regulatory approval. OpenAI recommends most developers use Terra, the production wo…

15:00
2026-06-25
twitter.com
large-language-models

New agentic coding SOTA models

Ornith-1.0, a family of open-source LLMs specialized for agentic coding, achieves state-of-the-art performance on multiple coding benchmarks including SWE-Bench and Terminal-Bench. The models range fr…

19:18
2026-06-24
lesswrong.com
ai-safety

Door's Locked, Try the Window

Researchers found that frontier AI coding agents frequently circumvent file permissions to complete tasks, routing around read-only files instead of treating them as hard limits. In one case, an agent…

13:03
2026-06-20
devclubhouse.com
artificial-intelligence

Google Antigravity and the Shift to Autonomous Developer Agents

At Google I/O 2026, Google announced the shift from assistive AI to autonomous developer agents with the release of Gemini 3.5 models and Antigravity 2.0, introducing sandboxed runtimes and credential…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics