cd/entity/AutomationBench· home› entities› AutomationBench
grep -l @automationbench /news/*.json | wc -l → 21

AutomationBench

mentions 21 type Organization page 1/2 feed RSS

// recent coverage 21 mentions

07:39
2026-10-07
dev.to
ai-research

My checklist for reading computer use benchmarks

A developer reviewed roughly 40 benchmark repositories and 230 sources to compile a five-question checklist for interpreting computer-use benchmark scores, finding that identical benchmark names can h…

21:31
2026-09-30
forgeeks.net
large-language-models

Gemini 4 Argon’s safety case now hinges on 1 million tokens

Google began rolling out Gemini 4 Argon on September 30, 2026 through its Fairwind Program to a selected group of trusted cyber defenders, with broader access undated, while claiming the model is its …

09:44
2026-09-28
huggingface.co
ai-agents

Holo4: powering generalist computer-use agents

H Company released Holo4, a new series of generalist computer-use agentic models in two sizes — 27B dense and 35B-A3B Mixture of Experts — both available on the H Models API, alongside an updated Holo…

20:50
2026-09-22
mashable.com
artificial-intelligence

ChatGPT-6 Sol and Luna are here: Pricing and benchmarks data

OpenAI launched two new ChatGPT models, GPT-6 Sol and GPT-6 Luna, on Tuesday, hours after Anthropic released Claude Opus 5.5, according to an OpenAI blog post. OpenAI priced GPT-6 Sol at $2 per millio…

07:09
2026-09-03
byteiota.com
large-language-models

Claude Fable 5.1: The Cache Cut That Changes Agent Costs

Anthropic shipped Claude Fable 5.1 on September 1, cutting cache read prices by 75% from $1.00 to $0.25 per million tokens, which reduces the real cost of highly agentic workloads by up to 45%. The mo…

00:00
2026-09-02
mindstudio.ai
artificial-intelligence

Claude Opus 5.1 Benchmarks: How Much Better Is It Than Opus 5?

Anthropic's Claude Opus 5.1 model shows significant benchmark gains over its predecessor Opus 5, with scientific research capability roughly doubling to over 50% on Terminal Bench Science 1.0 and busi…

21:25
2026-09-01
blog.kilo.ai
large-language-models

Claude Fable 5.1 Is Live in Kilo

Anthropic released Claude Fable 5.1, now available in Kilo, with cache read prices cut 75% to $0.25 per million tokens and benchmark gains on Terminal-Bench 4.0 (42% to 55.8%) and AutomationBench (17.…

02:10
2026-08-28
byteiota.com
large-language-models

Gemini 3.7 Flash: What Developers Need to Know Now

Google shipped Gemini 3.7 Flash on August 13, introducing breaking API changes including the replacement of the integer `thinking_budget` with a string enum `thinking_level`, removal of sampling param…

07:08
2026-08-17
sourcefeed.dev
artificial-intelligence

Gemini 3.7 Flash Is a Land Grab for Agent Workloads

Google shipped Gemini 3.7 Flash on August 13, three weeks after 3.6 Flash, with benchmark gains of 43.6% vs 34.4% on FrontierCode 1.1 Main, 65.3% vs 49.0% on DeepSWE v1.1, and 30.4% vs 17.0% on Automa…

16:35
2026-08-14
dev.to
artificial-intelligence

Gemini 3.7 Flash Makes Agent Cost the Feature

Google released Gemini 3.7 Flash on August 13, three weeks after 3.6 Flash, with a pricing strategy that cuts costs by half until the end of 2026. The model shows significant improvements on coding be…

00:08
2026-08-14
sourcefeed.dev
artificial-intelligence

Gemini 3.7 Flash's Half-Price Launch Has an Expiration Date

Google shipped Gemini 3.7 Flash on August 13, three weeks after Gemini 3.6 Flash, with benchmark gains in agentic coding tasks such as DeepSWE v1.1 jumping from 49.0% to 65.3%. The launch price is $0.…

22:07
2026-07-31
dev.to
large-language-models

Claude Sonnet 5 vs Opus 5: A Real-World Comparison (2026)

Anthropic's Claude Sonnet 5 and Claude Opus 5, released in mid-2026, offer distinct strengths and pricing trade-offs. Sonnet 5 excels at high-volume coding and content tasks with a 72.7% SWE-bench Ver…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics