What i run with my Strix Halo
A user reports running large language models on a 128GB Bosgame M5 Strix Halo mini PC purchased used for 1800€ on eBay, achieving 30 tokens per second decode with Qwen 3.8 27B and 50 tokens per second…
A user reports running large language models on a 128GB Bosgame M5 Strix Halo mini PC purchased used for 1800€ on eBay, achieving 30 tokens per second decode with Qwen 3.8 27B and 50 tokens per second…
Ant Group's inclusionAI released Ling 3.0 Flash, an open-weights AI model that scores 38 points on the Artificial Analysis Intelligence Index, making it the smartest open model under 124 billion total…
Anthropic released four operational controls for Claude Managed Agents, including per-session spend caps, region-pinned inference at a 1.1x in-region rate, repository-loaded skills, and a declarative …
Vercel's AI Gateway now offers Ling 3.0 Tiny from ANT Group, a free-to-use model with 7.9B total parameters and about 1.3B active per token, a 256K token context window, and up to 32K output tokens, a…
AntGroup released Ling 3.0 Flash, a 124-billion-parameter hybrid-reasoning Mixture-of-Experts model with 5.1 billion active parameters per token, matching or beating its 1-trillion-parameter flagship …
Kilo announced that Ling 3.0 Flash, the latest model from inclusionAI (Ant Group), is now available on its platform and free for a limited time. The model features 124B total parameters with only ~5.1…
Ant Group's Ling 3.0 Flash, a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token, is now available on Vercel's AI Gateway free for three weeks through August 3rd. The …