cd /news/large-language-models/grok-4-7-nears-anthropic-s-models-at… · home topics large-language-models article
[ARTICLE · art-136936] src=vibeleaderboard.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Grok 4.7 nears Anthropic's models at half the cost on analysis work

XAI shipped Grok 4.7, trained with a longer reinforcement learning run and native training on its own agent harness, and Artificial Analysis places it just behind Anthropic's models on a due diligence benchmark at roughly half the per-task cost of Opus 5. The release lands as Xiaomi's open-weight MiMo-V2.6 Pro tops Artificial Analysis's Intelligence Index among open models, and a Carnegie Mellon study found Claude, GPT-5, and Gemini fabricated medical diagnoses in 18% of cases when asked about an image that was never provided.

read2 min views2 publishedSep 22, 2026

xAI shipped Grok 4.7 with a longer reinforcement learning run and native training on its own agent harness. Artificial Analysis puts it just behind Anthropic's models on a due diligence benchmark, at roughly half the per task cost of Opus 5. Read: xAI shipped Grok 4.7 with a longer reinforcement learning run and native training on its own agent harness. Artificial Analysis puts it just behind Anthropic's models on a due diligence benchmark, at roughly half the per task cost of Opus 5. Read: Xiaomi released MiMo-V2.6 Pro and Flash, open-weight omnimodal models trained via scaled reinforcement learning. Pro now leads open models on Artificial Analysis's Intelligence Index while undercutting rivals on price, with training code and RL environments open-sourced. Read: Cloudflare made Python Workers generally available after a two-year preview, running Pyodide compiled to WebAssembly inside V8's workerd runtime. Python now has fully supported, production-ready status alongside JavaScript, though threading and multiprocessing remain unsupported. Read: LangSmith's new judge feature checks every production trace instead of sampling, evaluating more criteria per trace without cost scaling linearly with volume, and can flag safety or security issues fast enough to trigger an automated response. Read: A Carnegie Mellon study found Claude, GPT-5, and Gemini fabricated medical diagnoses in 18% of cases when asked about an image that was never provided, inventing conditions based only on a patient's stated age, race, and gender. Read: A study of recursive language models, which solve subtasks in isolated contexts, finds they generalize better out of domain than standard chain of thought, because chain of thought can exploit shortcuts hidden in the full reasoning trace that isolation rules out. Read: Nathan Lambert's congressional briefing notes assess the US-China balance of power in open-weight models, distinguishing open source from open weight and explaining why Chinese labs have led open releases since roughly April 2025.

── more in #large-language-models 4 stories · sorted by recency
── more on @xai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grok-4-7-nears-anthr…] indexed:0 read:2min 2026-09-22 ·