{"slug": "grok-4-7-nears-anthropic-s-models-at-half-the-cost-on-analysis-work", "title": "Grok 4.7 nears Anthropic's models at half the cost on analysis work", "summary": "XAI shipped Grok 4.7, trained with a longer reinforcement learning run and native training on its own agent harness, and Artificial Analysis places it just behind Anthropic's models on a due diligence benchmark at roughly half the per-task cost of Opus 5. The release lands as Xiaomi's open-weight MiMo-V2.6 Pro tops Artificial Analysis's Intelligence Index among open models, and a Carnegie Mellon study found Claude, GPT-5, and Gemini fabricated medical diagnoses in 18% of cases when asked about an image that was never provided.", "body_md": "xAI shipped Grok 4.7 with a longer reinforcement learning run and native training on its own agent harness. Artificial Analysis puts it just behind Anthropic's models on a due diligence benchmark, at roughly half the per task cost of Opus 5.\nRead: xAI shipped Grok 4.7 with a longer reinforcement learning run and native training on its own agent harness. Artificial Analysis puts it just behind Anthropic's models on a due diligence benchmark, at roughly half the per task cost of Opus 5.\nRead: Xiaomi released MiMo-V2.6 Pro and Flash, open-weight omnimodal models trained via scaled reinforcement learning. Pro now leads open models on Artificial Analysis's Intelligence Index while undercutting rivals on price, with training code and RL environments open-sourced.\nRead: Cloudflare made Python Workers generally available after a two-year preview, running Pyodide compiled to WebAssembly inside V8's workerd runtime. Python now has fully supported, production-ready status alongside JavaScript, though threading and multiprocessing remain unsupported.\nRead: LangSmith's new judge feature checks every production trace instead of sampling, evaluating more criteria per trace without cost scaling linearly with volume, and can flag safety or security issues fast enough to trigger an automated response.\nRead: A Carnegie Mellon study found Claude, GPT-5, and Gemini fabricated medical diagnoses in 18% of cases when asked about an image that was never provided, inventing conditions based only on a patient's stated age, race, and gender.\nRead: A study of recursive language models, which solve subtasks in isolated contexts, finds they generalize better out of domain than standard chain of thought, because chain of thought can exploit shortcuts hidden in the full reasoning trace that isolation rules out.\nRead: Nathan Lambert's congressional briefing notes assess the US-China balance of power in open-weight models, distinguishing open source from open weight and explaining why Chinese labs have led open releases since roughly April 2025.", "url": "https://wpnews.pro/news/grok-4-7-nears-anthropic-s-models-at-half-the-cost-on-analysis-work", "canonical_source": "https://www.vibeleaderboard.ai/intel/brief/2026-09-22", "published_at": "2026-09-22 11:16:17+00:00", "updated_at": "2026-09-22 11:53:49.915397+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-products", "ai-agents", "ai-safety"], "entities": ["xAI", "Grok 4.7", "Anthropic", "Opus 5", "Artificial Analysis", "Xiaomi", "MiMo-V2.6 Pro", "Carnegie Mellon"], "alternates": {"html": "https://wpnews.pro/news/grok-4-7-nears-anthropic-s-models-at-half-the-cost-on-analysis-work", "markdown": "https://wpnews.pro/news/grok-4-7-nears-anthropic-s-models-at-half-the-cost-on-analysis-work.md", "text": "https://wpnews.pro/news/grok-4-7-nears-anthropic-s-models-at-half-the-cost-on-analysis-work.txt", "jsonld": "https://wpnews.pro/news/grok-4-7-nears-anthropic-s-models-at-half-the-cost-on-analysis-work.jsonld"}}