KV Cache Size Calculator
A developer released a Python tool that calculates and plots KV cache size versus context length from HuggingFace config.json files. The tool supports standard MHA/GQA, MLA, hybrid architectures, and …
A developer released a Python tool that calculates and plots KV cache size versus context length from HuggingFace config.json files. The tool supports standard MHA/GQA, MLA, hybrid architectures, and …
A developer has published a configuration guide for running the Qwen3.8-Flash-Next 125B MoE model with llama.cpp on an RTX 4090 24GB system, achieving up to 29 tokens per second decode speed. The setu…
Tencent released Hy4 preview, a 770B-parameter open-weight Mixture-of-Experts language model with 49B active parameters and a 1M token context window, under Apache 2.0. In a blind evaluation across 20…
Tencent's 770B-parameter Hy4 preview model, a Mixture-of-Experts architecture with 49B activated parameters per token, is now deployable locally via prebuilt Docker images for vLLM and SGLang, requiri…
Lantern, a new free service, lets artists and creators upload images to generate one-way fingerprints and scans publicly available AI training datasets on Hugging Face and ModelScope to notify them if…
Z.ai, formerly Zhipu AI, released GLM-5.3-Flash on August 26, its first natively multimodal model in the GLM-5 series, running entirely on domestically produced Chinese AI accelerators. The 320-billio…
Alibaba released Qwen3.8-Flash-Next, an open-weight multimodal model with a 125B-parameter Mixture-of-Experts backbone plus 51B n-gram embedding parameters and 6B active parameters per token, which ou…
Alibaba released three new AI models in August 2025, including the flagship Qwen3.8-Max with 2.4 trillion total parameters (95 billion active), the lightweight Qwen3.8-27B with 27 billion parameters, …
An engineer has built Timeline Studio, an open-source, local-first AI video editor that runs entirely in the browser, using WebGPU, WASM, and ONNX to process media and AI models locally instead of upl…
Alibaba's Qwen3.8-27B, a 27-billion-parameter dense model released around August 14, matches or beats Anthropic's Claude Opus 4.6 on coding benchmarks, scoring 61.7 on SWE-bench Pro versus Opus 4.6 Ma…
Alibaba's Qwen team released Qwen3.8-27B, an open-weight multimodal model with 27 billion parameters and a native 262,144-token context window, available under Apache 2.0 on Hugging Face and ModelScop…
MiniMax, the Shanghai-based AI company founded by Yan Junjie, released Music 3 on August 13, 2026, an open-weights model that generates complete songs of up to five minutes from lyrics and a productio…
NVIDIA released Nemotron 3.5 Lightning on 2026-08-11, a 31.6B-parameter mixture-of-experts model with ~3.6B active parameters per token, hybrid Mamba-Transformer architecture, multi-token prediction, …
Alibaba Cloud launched its M890 AI supernode in Ulanqab, Inner Mongolia, on August 12, 2026, enabling enterprises to provision 64-card high-speed-interconnect computing units for inference on mixture-…
Alibaba will release open weights for Qwen3.8-Max and Qwen3.8-27B this week on Hugging Face and ModelScope, marking the first time a Max-class model is open-sourced. The 2.4-trillion-parameter Mixture…
Tencent opened its Hy3 large language model to global users on August 5, 2026, making it available through WorkBuddy, Miora, and Tencent Cloud TokenHub, with free access until August 31, 2026. Hy3, a …
Alibaba released Qwen3.8-Max on Monday, its most capable AI model, with 2.4 trillion total parameters and 95 billion active, and will open-source the weights on Hugging Face and ModelScope next week, …
Alibaba Group Holding Ltd. shares rose as much as 6% in Hong Kong trading on Monday after the company unveiled Qwen3.8-Max, a 2.4 trillion parameter mixture-of-experts model that Alibaba claims trails…
Rescene, a free AI agent aggregator, has been released, requiring no API key. It aggregates 7 free providers and 18 model entries, featuring intelligent routing, browser automation, and Computer Use, …
Moonshot AI released the full Kimi K3 open weights on July 27, 2026, a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window and a 1.4 TB download. The model uses nat…