cd /news/artificial-intelligence/benchmark-headlines-break-down-cost-… · home topics artificial-intelligence article
[ARTICLE · art-104597] src=vibeleaderboard.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Benchmark headlines break down: cost per task, not per token, decides the day

Artificial Analysis reported that Qwen3.8 Max's per-token price fell while its cost per Intelligence Index task more than doubled versus Qwen3.7 Max, and its GDPval-AA Elo lead over Kimi K3 came from taking roughly four times more turns per task, highlighting that per-token pricing misleads. The same model regressed ten points on AA-Omniscience (+14 to +4) with hallucination rate roughly doubling at flat accuracy, and Artificial Analysis patched the Intelligence Index to v4.1.1 with changed graders, making cross-version comparisons invalid. Meta's Muse Spark 1.2 landed on the cost-per-task Pareto frontier, and Kimi K3 reached GA in Copilot before a rollout pause, while Agent Plugins shipped as a cross-client standard from OpenAI and Cursor.

read2 min views1 publishedAug 7, 2026

The day's most useful number was not a leaderboard rank but a denominator. Artificial Analysis showed Qwen3.8 Max cutting per-token price while more than doubling cost per Intelligence Index task, winning GDPval-AA Elo largely by taking roughly four times as many turns, and regressing ten points on AA-Omniscience as it stopped abstaining — three different ways a headline score can hide latency, spend and hallucination risk, arriving alongside an Intelligence Index grader change that makes cross-version comparisons invalid. Against that, Meta's Muse Spark 1.2 landed on the cost-per-task frontier and Kimi K3 reached GA in Copilot before a rollout , giving builders cheaper defaults whose real economics still need per-workload measurement. Underneath the model news, the plumbing converged: Agent Plugins shipped as a cross-client standard from OpenAI and Cursor on the same day the MCP ecosystem moved toward stateless servers and web-side tool interfaces. Debate: Qwen3.8 Max is the cleanest case yet that per-token pricing misleads: cost per task more than doubled versus Qwen3.7 Max even as token rates fell, and its GDPval-AA Elo lead over Kimi K3 came from taking about four times more turns per task. Watch: The same model regressed ten points on AA-Omniscience (+14 to +4) with hallucination rate roughly doubling at flat accuracy — a model that stopped abstaining is materially riskier in unattended agent loops than its index score suggests. Method: Artificial Analysis patched the Intelligence Index to v4.1.1 with changed graders, so scores cited across index versions are no longer directly comparable — check the version before quoting a number in a design doc. Release: Cheaper defaults arrived from two directions: Muse Spark 1.2 puts Meta on the cost-per-task Pareto frontier, while open-weight Kimi K3 hit GA in Copilot and is callable through the LangSmith gateway — though GitHub has d the K3 rollout while it mitigates an incident. Tooling: Agent Plugins landed as an open standard from OpenAI Devs and Cursor on the same day, letting a skill or MCP configuration authored once ship to Codex, Cursor, Copilot and VS Code without per-client repackaging. Watch: The agent-facing web stack advanced in parallel — a stateless next-generation MCP that runs as ordinary request handlers, WebMCP interfaces for existing sites, an agent-first browser in V8 isolates, managed retrieval via Cloudflare AI Search, and a protocol sketch for a readable, callable, payable agentic internet. Method: Two papers push back on summarization as a default context strategy: FinPerMA finds summarization drops the preference signal personalization depends on, with plain retrieval outperforming it, while SONAR argues code summaries for agent consumption should optimize correctness and abstraction level rather than polish. Release: OpenAI moved GPT-5.6 Sol behind all paid chats with a quantified factuality gain, expanded GPT-5.6 Luna to free users and removed free-tier chat limits, changing what you can assume about the model your users are actually on.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @artificial analysis 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/benchmark-headlines-…] indexed:0 read:2min 2026-08-07 ·