Most Neoclouds Suck At Security
SemiAnalysis's ClusterMAX 3.0 testing found that most neoclouds have serious security vulnerabilities, with five frightening patterns identified. The company urges neocloud operators and users to upda…
SemiAnalysis's ClusterMAX 3.0 testing found that most neoclouds have serious security vulnerabilities, with five frightening patterns identified. The company urges neocloud operators and users to upda…
Chinese AI models are closing the performance gap with US rivals while undercutting them on price, forcing Silicon Valley to cut prices, according to users and analysts. Zhipu AI's GLM-5.2 costs US$4.…
Tencent released Hy4 preview on August 28, a 770B parameter open-weight model under Apache 2.0, with only 49B active parameters per token and a 1 million token context window, targeting software engin…
Two major labs released dense ~30B parameter multimodal models with Apache 2.0 licenses in August 2026: Meta's Muse Glimmer 30B and Alibaba's Qwen3.8-27B. Both are designed for local GPU inference on …
A new census of 54 AI models across 13 vendors, published August 30, 2026, reveals that maximum output token limits—the amount a model can write in a single response—are rarely disclosed, especially i…
A developer's benchmark of 22 GPT-5.6 Sol configurations found that High reasoning effort in Fast mode offers the best performance-cost-speed trade-off, scoring 57 on the Artificial Analysis Intellige…
Qwen 3.8 27B, running locally on a 16GB RAM MacBook Pro, outperformed OpenAI's GPT 5.6 Luna Max on the DABstep benchmark at over 17 times lower cost, with electricity costs under $0.50 versus over $8.…
Tencent released Hy4 preview, a 770-billion-parameter open-source flagship model with 49 billion active parameters and a one-million-token context window, targeting long-horizon software engineering, …
Beijing-based Moonshot AI released its Kimi K3 model with open weights, positioning China's open software approach as safer than the closed strategies of Silicon Valley rivals OpenAI and Anthropic. Th…
WARP, an embeddable inference engine written in C, now runs the full 2.78-trillion-parameter Kimi K3 model on a 64 GB MacBook Pro at about 0.6 tokens per second, and the 313-billion-parameter GLM-5.3-…
Nvidia is paying $6 billion to license AI startup Poolside's models, aiming to build one of the world's most powerful open-weight AI models to compete with Chinese rivals like DeepSeek and Kimi K3, ac…
Tencent claims its Hunyuan 3.0 foundation model, with 295 billion parameters and deep integration into WeChat and its Yuanbao AI assistant, outperforms rivals in internal tests, positioning it against…
Z.ai's stealth preview model, branded 'ox-alpha' on OpenRouter.AI and OpenCode.AI, has launched as GLM-5.3-Flash, an efficient open-weights model priced at $0.07 per million input tokens and $0.25 per…
Apple's upcoming M5 Ultra, with 512GB memory and 1.2TB/s bandwidth, is projected to cost $15,000–17,000 when it ships in October, making local AI inference for models like DeepSeek V4 Pro, Kimi K3, an…
ZAI's GLM-5.3-Flash, reportedly code-named Ox Alpha, drew significant online hype before release, but an independent SimpleBench score fell short of Google's Gemini models, undercutting claims it was …
Tencent released Hy4 preview, a 770B-parameter open-weight Mixture-of-Experts language model with 49B active parameters and a 1M token context window, under Apache 2.0. In a blind evaluation across 20…
Telnyx has launched GLM-5.3, a 753B-parameter reasoning model with a 1M token context window, on its Inference API, claiming it matches Kimi K3 on the Artificial Analysis Intelligence Index at roughly…
Only two AI labs in 2025–2026 published a dated, on-record reason before withholding open-weights releases, according to a new analysis: Z.ai held GLM-5.3 for fourteen days in August 2026, and OpenAI …
Four reproducible parser failures in vLLM 0.26.0, 0.27.1, and 0.28.0 can drop or garble tool calls and reasoning content while still returning HTTP 200, according to the Ingot team's tests on CPU with…
Chinese open-weight AI models are gaining traction with U.S. enterprises, with Ramp's AI Index showing the share of businesses paying for model serving platforms rose to 6.1% in July 2026 from 4.5% in…