Kog says GPUs aren't the wrong fit
Kog, an eleven-person French startup, claims that the next leap in AI inference speed will come from software optimizing existing GPUs rather than new silicon, and has demonstrated 30× faster decoding…
Kog, an eleven-person French startup, claims that the next leap in AI inference speed will come from software optimizing existing GPUs rather than new silicon, and has demonstrated 30× faster decoding…
Caveman 2, an MIT-licensed add-on for coding agents including Claude Code, Codex, Gemini, Cursor, and Windsurf, shipped version 1.10 this week, cutting input tokens by 33.2% via a new local proxy and …
Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter multimodal open model under Apache 2.0, designed for local agent workloads on consumer hardware. The model scores 76 on SWE-Be…
Two independent benchmark sites, BenchLM.ai and llm-stats.com, compared DeepSeek's V4 Flash 0731 update with Alibaba's Qwen3.6-27B, finding that DeepSeek V4 Flash is roughly 12× cheaper per token and …
Chinese API resellers are selling Claude and Codex tokens at up to 90% discounts, or 10% of official prices, according to a Show HN post on 2 August by xiaoxumz11 and a May ChinaTalk analysis by Zilan…
A developer writing as Rituraj on Dev Genius in late July tested three retrieval methods on a corpus of paired current and deprecated documents, finding that vector RAG returned both versions with cos…
A community tester ran DeepSeek-V4-Flash-0731 on a Bosgame M5 mini PC with an RTX PRO 6000 Max-Q eGPU, achieving 44 to 60 tokens per second depending on quantisation, with no hyperscaler involved. The…
DeepSeek's V4 Flash model graduated from preview to official public-beta release on 31 July 2026, with post-training improvements that boosted its Toolathlon score from 51.8 to 70.3 on DeepSeek's own …
OpenAI open-sourced its Codex Security CLI, a command-line tool that scans code for vulnerabilities and suggests patches, under the Apache 2.0 license on Wednesday. The tool, previously codenamed Aard…
Agenta, an MIT-licensed open-source workspace for building and running AI agents, has shipped a self-hosted alternative to Anthropic's Claude Cowork that supports background execution, shared workspac…
Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, nearly quadrupling the previous record of 7.8% set by OpenAI's GPT-5.6 Sol, according to the ARC Prize team. The result marks the mos…
Amazon Web Services launched Anthropic's Claude Opus 5 on Amazon Bedrock on 24 July 2026, pricing the model at $5 per million input tokens and $25 per million output tokens — half the cost of the flag…
AMD is investing up to $5 billion in Anthropic, the AI lab behind Claude, in a deal that includes Anthropic deploying up to two gigawatts of AMD's top-end AI chips starting in the first half of 2027, …
Alibaba's Qwen3.6 35B A3B scores 32 on Artificial Analysis's Intelligence Index v4.1, outperforming Google's Gemma 4 26B A4B at 26, with Qwen winning 18 of 22 evaluations tested. However, Gemma 4 is 6…
Alibaba previewed Qwen 3.8, an open-weight model with over one trillion parameters that handles images, video, documents and text, claiming it is second only to Anthropic's Fable 5. The Qwen team says…
NVIDIA announced at SIGGRAPH 2026 that six creative tools — Adobe, Affinity by Canva, Blender, Boris FX, Foundry, SideFX and Epic's Unreal Engine — have adopted the Model Context Protocol (MCP), enabl…
Anthropic announced on Friday that Claude Fable 5 will remain in Max and Team Premium plans from 20 July but at half the regular usage limits, with Pro and Team Standard subscribers losing access and …
A March 2026 study from MIT's Computer Science and Artificial Intelligence Laboratory analyzing 809 large language models released between 2022 and 2025 found that 80 to 90% of frontier AI model perfo…
Anthropic confidentially filed for an IPO on 1 June, with Polymarket putting the chance of a 2026 listing at 65%, and investor Gavin Baker estimating on the All-In podcast that Anthropic would trade a…
NVIDIA published a new CPU category on 7 July with Vera, an Arm server CPU designed to address the agentic-AI bottleneck by maximizing single-threaded performance rather than core density. In testing …