Evolution of Attention over the Years
A technical survey traces five rewrites of transformer attention since 2017, framing every variant as either storing less or reading less. It details how multi-query and grouped-query attention shrink…
A technical survey traces five rewrites of transformer attention since 2017, framing every variant as either storing less or reading less. It details how multi-query and grouped-query attention shrink…
Anthropic's claim that Moonshot and DeepSeek route user requests to Anthropic models is false, according to a rebuttal that attributes the traffic to Chinese gray-market "transit station" AI gateways.…
InclusionAI is offering free API access to its Ling-3.0-Flash-VL multimodal model and releasing the weights under the MIT license, the company said. The model uses a mixture-of-experts design with 124…
A deep-research pass conducted through four LLMs — ChatGPT, Gemini, Grok, and GLM — concluded that TRIZ should not be applied to software debugging until a failure's cause is established through evide…
Starkomand launched v1.0.2, a macOS Apple silicon workspace that runs Claude Code, Codex, Cursor, Grok, GLM and other CLI coding agents side by side on a single infinite canvas with local dictation an…
A laid-off engineering executive built Bakeoff, a private evaluation that runs real bugs from his own codebases through the latest OpenAI and Anthropic models in their own harnesses plus GLM in pi, sc…
Chinese AI models now dominate OpenRouter's weekly usage rankings, with 8 of the top 10 models by token usage being Chinese, including Tencent's Hy4 Preview at #1 with 20.4 trillion tokens. These eigh…
OpenRouter's multi-provider routing causes significant performance variability for the same AI model, according to a developer who runs the iMessage assistant Olly and has processed over 18 million me…
OpenAI's GPT-6 Astra, a long-running agentic AI model, can execute multi-step administrative tasks across websites and tools without step-by-step prompting, as demonstrated in a hands-on test involvin…
Sequoia Capital told 80 portfolio founders that owning AI models down to the weights is now a performance edge, not a sacrifice, according to a talk by partner Sonya Huang titled 'Own Your Intelligenc…
Developer siropkin released chrome-bridge, an open-source tool that lets any AI agent drive a user's real, logged-in Chrome browser via one unpacked extension and a zero-dependency Node CLI, requiring…
Philips Hue announced Custom AI Behaviors at IFA, an AI assistant feature that lets users create complex light automations using natural language, with the automation engine running locally on the new…
A new analysis questions whether large language models can be trusted to build secure codebases, highlighting the widening performance gap between frontier models and the persistent risks of AI-genera…
Qwen, DeepSeek, and GLM are closing the gap with trillion-parameter models by focusing on post-training optimization and extended reasoning rather than raw parameter count, according to an analysis on…
Broadcom announced at VMware Explore in Las Vegas a suite of private AI services, including an AI gateway for model governance, a partnership with MetalSoft to reduce bare-metal provisioning from week…
In August, Chinese AI labs released a series of frontier and near-frontier open-weights models, including DeepSeek-v4-Flash-0731, Alibaba's Qwen3.8-27B, Z.ai's GLM 5.3 and GLM 5.3 Flash, and Qwen3.8-F…
A developer released a Python tool that calculates and plots KV cache size versus context length from HuggingFace config.json files. The tool supports standard MHA/GQA, MLA, hybrid architectures, and …
Tencent's Hy Team released Hy4 preview, a 770B-parameter Mixture-of-Experts flagship model with 49B activated parameters per token, featuring Gated DeepSeek Sparse Attention and a 1M context length. T…
Open-source and free AI models are becoming nearly as capable as frontier American models, with local models like Ornith 1.5 9B running on consumer hardware and free API models from Nvidia and others …
A systematic comparison of Chinese LLM tool-calling compatibility as of August 2026 finds that while all five major families (DeepSeek, GLM, Qwen, Kimi, MiniMax) expose OpenAI-style endpoints, signifi…