{"slug": "qwen-s-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5-s-prose", "title": "Qwen's agentic crown lasted hours — and r/ClaudeAI is fighting Opus 5's prose", "summary": "Qwen3.8 Max briefly led Artificial Analysis' Agentic Index by 0.1 points before a benchmark version bump moved every score, with Claude Opus 5 now topping the chart at 59.2 versus Qwen3.8 Max's 58.4. Meanwhile, r/ClaudeAI's top thread criticizes Opus 5 for verbose, jargon-heavy documentation, and a 40,000-run experiment found humans missed 1 in 3 threats when approving agent commands.", "body_md": "# Qwen's agentic crown lasted hours — and r/ClaudeAI is fighting Opus 5's prose\n\nQwen3.8 Max led the agentic index by 0.1 points for a few hours before a version bump moved every score. Plus: the Opus 5 prose revolt, two measurements of how review fails, and the skills wave.\n\nThe most-upvoted AI story of the day says Qwen3.8 Max just took the agentic crown from Claude Opus 5. It did — by 0.1 points, for a few hours, until a benchmark version bump moved every score on the chart. The timeline is better than the ranking. Meanwhile r/ClaudeAI's top thread is a revolt against how Opus 5 writes, a 40,000-run experiment puts a number on how badly humans review their agent's commands, and the skills wave still owns GitHub trending. The daily pulse of AI coding tools — what shipped, what matters, what's next.\n\n**In this issue**\n\n•\n\n[Qwen's agentic lead lasted hours; Opus 5 tops it 59.2](https://ai-news.ghost.io/qwens-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5s-prose/#qwen38-max-led-the-agentic-index-by-01-for-hours-%E2%80%94-then-a-version-bump-moved-every-score)\n\n•\n\n[r/ClaudeAI says Opus 5 writes docs nobody can read](https://ai-news.ghost.io/qwens-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5s-prose/#rclaudeais-top-thread-says-opus-5-writes-docs-nobody-can-read)\n\n•\n\n[Humans missed 1 in 3 threats approving agent commands](https://ai-news.ghost.io/qwens-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5s-prose/#humans-missed-1-in-3-threats-in-a-40000-run-agent-approval-game)\n\n•\n\n[54% of AI-written security patches failed or added flaws](https://ai-news.ghost.io/qwens-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5s-prose/#off-by-1-labs-found-539-of-ai-written-security-patches-failed-or-added-flaws)\n\n•\n\n[DeepSeek warns of a big price rise in its own docs](https://ai-news.ghost.io/qwens-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5s-prose/#deepseek-warns-of-a-%E2%80%9Csignificant%E2%80%9D-price-increase-in-its-own-pricing-docs)\n\n•\n\n[Trending is all skills, and installs were the trust signal](https://ai-news.ghost.io/qwens-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5s-prose/#everyones-top-of-trending-is-the-same-word-skills)\n\n## The leaderboard everyone quoted today\n\n### Qwen3.8 Max led the agentic index by 0.1 for hours — then a version bump moved every score\n\nYesterday's biggest story on Hacker News and r/LocalLLaMA said Qwen3.8 Max had taken the top spot on Artificial Analysis' Agentic Index. It had — by 0.1 points, briefly. Archived snapshots pin the timeline. August 5: Opus 5 leads at 55.3, Qwen3.8 Max not yet evaluated. August 6 morning: Qwen3.8 Max debuts at 55.4. Same day, Artificial Analysis ships Intelligence Index v4.1.1, re-versioning τ³-Banking (half the agentic score) and upgrading graders; the live chart now reads Opus 5 **59.2**, Qwen3.8 Max 58.4. The “best overall model” framing was never right either: the Agentic Index averages two benchmarks, and the nine-benchmark Intelligence Index has Opus 5 first at 63 with Qwen3.8 Max outside the top ten. A 0.1-point lead on a sub-index was the day's top story. Check the chart's date before repeating its rank.\n\n[Qwen 3.8 Max now ranked as best overall model ahead of Opus 5 by Artificial Analysis agentic index]\n\nby[u/anderspitman]in[LocalLLaMA]\n\n## What developers are saying about the model they just got\n\n### r/ClaudeAI's top thread says Opus 5 writes docs nobody can read\n\nThe thread asks how to stop Opus 5 writing documentation. The sub's bot summary after 80 comments reports consensus: verbose, jargon-heavy, ignoring CLAUDE.md, skills and memories. The top reply names the vocabulary — “load bearing”, “seam”, “blast radius”, “surface”. The fixes: plan with Fable, delegate code to Sonnet 5; enforce ASD-STE100 simplified technical English; add a stop-hook rejecting those words. We banned this register in our style rules yesterday. It's real. If your docs run on Opus 5, add that hook.\n\n[Opus 5 is literally useless for documentation]\n\nby[u/Sneaky_Tangerine]in[ClaudeAI]\n\n## Reviewing what your agent writes and runs\n\n### Humans missed 1 in 3 threats in a 40,000-run agent-approval game\n\nScale X ran a browser game where you approve agent commands. Across 40,000 runs, mean accuracy was **66.3%**. Miss rates split by type: 11.7% for destructive commands like rm -rf /, but 33.4% for exfiltration and 35.0% for reading ~/.aws/credentials. Most-approved threat: npm run analyze at 64.7%, since the danger is whatever package.json points at. The author's caveat: 34% of commands shown were threats, under time pressure. That's not a real day. Assume your approval step catches only the loud attacks.\n\n### Off-by-1 Labs found 53.9% of AI-written security patches failed or added flaws\n\nOff-by-1 Labs at 1Password had ChatGPT-5.5 and Opus 4.8 generate **6,080** patches for six recently disclosed vulnerabilities, including a Gemini CLI remote-code-execution bug and a Chrome use-after-free. Only 26.0% fully fixed the flaw without changing behavior. 20.1% fixed it but changed behavior. 53.9% failed, introduced a new vulnerability, or both. A third of the successful patches were narrow input checks, not root-cause fixes, so the bug returns if the code becomes reachable elsewhere. Never merge a model's security patch unread.\n\n## What it costs\n\n### DeepSeek warns of a “significant” price increase in its own pricing docs\n\nDeepSeek's pricing page now carries a line the coverage paraphrased: “We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected.” There's no number and no date. The specific plan is “subject to official notice.” So price what you have today: deepseek-v4-flash runs $0.14 per million input tokens on a cache miss and $0.28 output, deepseek-v4-pro $0.435 and $0.87. If DeepSeek's floor is what your cost model rests on, re-run the numbers this week.\n\n## 5 trending AI coding tools right now (and the problem no list mentions)\n\n### Everyone's top-of-trending is the same word: skills\n\nOpen GitHub trending any day this week and it's wall-to-wall agent skills — packaged instruction sets your coding agent loads on demand. The big three:\n\n1. [obra/superpowers](https://github.com/obra/superpowers?ref=ai-news.ghost.io) (268K⭐) — the skills framework the others orbit. “Often copied, never beaten” per its fans; “does out-of-the-box Fable make it obsolete?” per its skeptics. That argument is itself a sign it won.\n\n2. [mattpocock/skills](https://github.com/mattpocock/skills?ref=ai-news.ghost.io) (208K⭐) — “Skills for Real Engineers,” straight from Pocock's own .agents directory. Six months old.\n\n3. [addyosmani/agent-skills](https://github.com/addyosmani/agent-skills?ref=ai-news.ghost.io) (83K⭐) — production-grade skills from Chrome's engineering leadership, trending daily as of this morning.\n\nGoogle itself is now on the board: [google/skills](https://github.com/google/skills?ref=ai-news.ghost.io) (16K⭐, “Agent Skills for Google products and technologies”) is trending today. When the platform vendors start shipping into an ecosystem, the ecosystem stops being a hobby.\n\n### The long tail is where it gets interesting\n\n4. The niche skill packs — this week's daily/weekly lists include [book-to-skill](https://github.com/virgiliojr94/book-to-skill?ref=ai-news.ghost.io) (turn any technical-book PDF into a skill, 18K⭐), [reverse-skill](https://github.com/zhaoxuya520/reverse-skill?ref=ai-news.ghost.io) (authorized-pentest skill router, 20K⭐), and [i-have-adhd](https://github.com/ayghri/i-have-adhd?ref=ai-news.ghost.io) (18K⭐) — a skill whose whole job is stopping your agent from burying the answer. When the long tail gets this specific, an ecosystem has arrived.\n\n### The non-skills entry\n\n5. [antirez/ds4](https://github.com/antirez/ds4?ref=ai-news.ghost.io) (21K⭐) — Salvatore Sanfilippo's from-scratch local inference engine for DeepSeek 4 Flash (Metal/CUDA/ROCm). The Redis author writing a bare-metal engine for last week's hottest open model is the most “2026” sentence available.\n\nBubbling under: [cloudflare/computer](https://github.com/cloudflare/computer?ref=ai-news.ghost.io) (“give your agent a computer,” 5.2K⭐ and still climbing), [code-review-graph](https://github.com/tirth8205/code-review-graph?ref=ai-news.ghost.io) (local-first code intelligence for MCP — the plug standard AI coding agents use — 29K⭐).\n\n### The npm scoreboard\n\nWeekly downloads, week ending Aug 6: @openai/codex 16.4M · @anthropic-ai/claude-code 11.4M · opencode-ai 2.2M. Codex still leads by about 5M. But Claude Code grew faster last week, 11.2% against Codex's 6.4%, and the gap narrowed slightly. One week is not a trend; check it again next Friday.\n\n### And here's what none of the “top tools” posts mention\n\nYesterday at Black Hat, [Zenity Labs disclosed a credential-stealing campaign](https://labs.zenity.io/post/attackers-target-agents-via-the-skill-supply-chain?ref=ai-news.ghost.io) that ran through Vercel's skills.sh marketplace — a typosquatted skill family (fake Paperclip / Browser Use skills) that amassed **1.7M aggregate installs** (their caveat: installs, not unique users).\n\nThe mechanism is the story: the skills were clean copies of legitimate skills while they accumulated installs — then the content behind them was swapped to make agents download and run attacker code, harvesting SSH keys, cloud creds, git tokens, and CI configs. A time-of-check/time-of-use hole in how marketplaces vet.\n\nRead that against the list above: every trending roundup — including this one — ranks by stars and installs. **The install count was the trust signal being gamed.** Vercel and GitHub pulled everything within 12 hours of notification; more than 30% of the other dangerous skills Zenity found abuse Claude Code or OpenClaw as malware droppers (per Zenity's press release; the campaign write-up doesn't carry that stat); one reinstalls itself if deleted. And the detail the full report adds: PyPI caught this same actor's packages twice, each within about 2 hours — the skills marketplace let the trojanized family trend for about 3 weeks.\n\nTrend responsibly.\n\n## Also worth your time\n\n• [A Fable 5 agent with a domain and a $90 budget it can't spend without approval](https://www.reddit.com/r/ClaudeAI/comments/1vhp54h/i_gave_a_claude_fable_5_agent_a_domain_and_90_it/?ref=ai-news.ghost.io) — it named itself Cairn and keeps a blog. r/ClaudeAI, 231 points.\n\n• [A reported hidden “instantaneous” rate limit beyond the 5-hour and weekly ones](https://www.reddit.com/r/ClaudeAI/comments/1vhphmc/warning_hidden_instantaneous_plan_limit_not_just/?ref=ai-news.ghost.io) — unverified, actionable if it holds up. r/ClaudeAI.\n\n• [vLLM's serving stack ported to C++20](https://www.reddit.com/r/LocalLLaMA/comments/1vh9lx4/i_ported_vllms_serving_stack_to_c20_66_mib_binary/?ref=ai-news.ghost.io) — a 66 MiB binary, no Python at inference, output checked token-for-token. r/LocalLLaMA, 287 points.\n\n• [“Software development with AI is starting to feel like cooking steak”](https://blog.sydorets.com/en/posts/almost-no-skill-required-to-cook-a-steak/?ref=ai-news.ghost.io) — the essay behind a 404-comment Hacker News thread.\n\nKnow someone who'd want this in their inbox? Forward it — that's how this grows. And if we got something wrong, or you think we buried the real story today, hit reply. A person reads every one.\n\n*The New Way is human-curated — a person picks every story. The summaries are written with AI (Claude) and reviewed before we hit send.*", "url": "https://wpnews.pro/news/qwen-s-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5-s-prose", "canonical_source": "https://ai-news.ghost.io/qwens-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5s-prose/", "published_at": "2026-08-07 15:54:54+00:00", "updated_at": "2026-08-26 06:43:22.416054+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents"], "entities": ["Qwen3.8 Max", "Artificial Analysis", "Claude Opus 5", "r/ClaudeAI", "DeepSeek"], "alternates": {"html": "https://wpnews.pro/news/qwen-s-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5-s-prose", "markdown": "https://wpnews.pro/news/qwen-s-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5-s-prose.md", "text": "https://wpnews.pro/news/qwen-s-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5-s-prose.txt", "jsonld": "https://wpnews.pro/news/qwen-s-agentic-crown-lasted-hours-and-r-claudeai-is-fighting-opus-5-s-prose.jsonld"}}