cd/entity/DeepSWE· home entities DeepSWE
grep -l @deepswe /news/*.json | wc -l → 55

DeepSWE

mentions 55 type Organization page 2/3 feed RSS

// recent coverage 55 mentions

12:56
2026-08-13
deepswe.datacurve.ai
artificial-intelligence

Grok 4.6 /medium outperforms /high effort on DeepSWE

DeepSWE, a new long-horizon software engineering benchmark, reports that Grok 4.6 /medium outperforms /high effort on its leaderboard, which measures frontier coding agents on original tasks across 91…

01:08
2026-08-13
sourcefeed.dev
artificial-intelligence

DeepSeek V4 Pro Goes GA — Mind the Weights Gap

DeepSeek released the production version of its flagship model, DeepSeek-V4-Pro-0813, on August 13, ending a preview that began April 24, with the API alias now resolving to the new build. The model s…

06:09
2026-08-10
byteiota.com
artificial-intelligence

DeepSeek V4 Flash 0731: Benchmarks, Pricing, and Dev Setup Guide

DeepSeek released V4-Flash-0731 on July 31, 2025, and the smaller model now outperforms the larger V4-Pro-Preview across all nine published agentic benchmarks, including Terminal Bench 2.1 (82.7 vs. 7…

12:09
2026-08-03
byteiota.com
artificial-intelligence

DeepSeek V4 Flash 0731: Smaller Model, Better Agent Scores

DeepSeek released the V4-Flash-0731 model on July 31, a retrained 13-billion-active-parameter Mixture-of-Experts model that outperforms its own V4-Pro-Preview on all nine agent benchmarks, including a…

07:19
2026-07-31
technode.com
artificial-intelligence

DeepSeek puts V4-Flash API into public beta

DeepSeek has launched the public beta of its V4-Flash API, featuring upgrades for agent tasks, scoring 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. The update adds Responses API support and Codex a…

00:00
2026-07-26
together.ai
artificial-intelligence

Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Kimi K3 edges GPT-5.6 Sol on pass@4 (89.4% vs 85.8%) and costs 64% less per rollout ($4.65 vs $8.37), but GPT-5.6 Sol leads on pass@1 (72.7% vs 68.5%) and reliability (61 tasks solved on all four trie…

06:00
2026-07-25
dev.to
artificial-intelligence

Google mise sur des Gemini moins chers, pas plus forts

Google released three new Gemini models on Tuesday, emphasizing cost reduction over raw performance. The Gemini 3.6 Flash model uses 17% fewer output tokens than its predecessor, lowering prices, whil…

00:00
2026-07-24
together.ai
artificial-intelligence

Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

Kimi K3 matches Claude Fable 5 on DeepSWE benchmark quality with a 68.5% pass@1 versus Fable's 69.9%, but costs $4.65 per rollout compared to Fable's $13.41, delivering 2.8x more solved tasks per doll…

18:16
2026-07-21
byteiota.com
artificial-intelligence

Gemini 3.6 Flash: 17% Fewer Tokens, Live in Copilot

Google shipped Gemini 3.6 Flash on July 21, 2026, offering 17% fewer output tokens than its predecessor and available immediately in GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise users. …

16:44
2026-07-20
i-programmer.info
ai-agents

DeepSWE - Best Benchmark For Evaluating AI Coding Agents?

DeepSWE, a new benchmarking platform for AI coding agents, achieves clearer separation between frontier models than older benchmarks like SWE-bench by using contamination-free tasks across 91 reposito…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics