Qwen3.8-Flash-Next: How to Run Locally
Qwen released Qwen3.8-Flash-Next, a 125B-parameter open-weight multimodal MoE model built on the Qwen4 architecture with a 262K context window, which outperforms Claude-4.6-Opus (Max) and can run loca…
Qwen released Qwen3.8-Flash-Next, a 125B-parameter open-weight multimodal MoE model built on the Qwen4 architecture with a 262K context window, which outperforms Claude-4.6-Opus (Max) and can run loca…
Apple's M5 Ultra Mac Studio, announced August 25, 2026, offers up to 512GB of unified memory at 1.2TB/s in a single box, competing with clusters of four NVIDIA DGX Sparks or four AMD Ryzen AI Halo box…
Perplexity launched 'Portable Computer,' a local-first AI agent developed with NVIDIA, which runs entirely on-device using small models such as Qwen 3.8 and NVIDIA Nemotron 3.5 Lightning, ensuring use…
A developer debugging a knowledge graph pipeline on an NVIDIA DGX Spark found that a 49-billion-parameter open-weight reasoning model was silently generating long internal monologues, causing apparent…
Perplexity has released Portable Computer, a local-first build of its agentic Computer platform that runs on NVIDIA DGX Spark and RTX GPUs with 24 GB of VRAM, packaging the local model, inference engi…
Alibaba's Qwen 3.8 27B open-weight vision model, released under Apache 2.0, takes 21 minutes and 22,276 reasoning tokens to draw a pelican riding a bicycle as an SVG, according to Simon Willison's tes…
A developer has released a CUDA-optimized fork of antirez's h3.c for NVIDIA DGX Spark, achieving a ~15.5x speedup, while also providing a native Apple Silicon implementation for MiniMax-H3 inference w…
Alibaba's Qwen team released Qwen 3.8 27B, a dense 27 billion parameter multimodal model that scored 51 on the Artificial Analysis agentic index, trailing only Kimi K2, a 2.8 trillion parameter model.…
Alibaba's Qwen 3 27B, a dense open-weight multimodal model, can be run locally using DeepSeek Harness, achieving agentic performance close to Claude 4.5 on benchmarks. On a two-node NVIDIA DGX Spark c…
A developer's ongoing 32-week project to understand AI inference reached Phase 2, focusing on tensors, the core data structure of machine learning models. The developer demonstrated creating tensors i…
Alibaba's Qwen research lab released Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM, which defaults to an 'xhigh' reasoning effort that causes it to overthink simple tasks, taking…
An engineer rebuilt the execution backbone of their development workflow onto local large language models (LLMs) to eliminate recurring cloud costs. By purchasing an NVIDIA DGX Spark and combining it …
Simon Willison released CORS Chat, a web tool built with GPT-5.6-Sol xhigh, to test Qwen 3.8 27B running in LM Studio on his M5 MacBook Pro and an NVIDIA DGX Spark. The tool provides a web UI for Open…
StorageReview's lab-tested leaderboard for best local AI desktops in 2026 ranks the NVIDIA DGX Spark as the best overall deskside AI system, featuring a GB10 Grace Blackwell superchip with 128GB unifi…
Luisuantech's GP Spark enclosure adds up to 16TB of NVMe storage to the NVIDIA DGX Spark via a 100GbE connection, addressing the DGX Spark's lack of full-size M.2 slots. The 0.58-liter, four-bay box s…
LTX released LTX-2.5, an open weights world model for video generation optimized for local inference on NVIDIA RTX GPUs and NVIDIA DGX Spark, cutting VRAM requirements so creators can run a frontier m…
MSI announced that its XpertStation WS300 and EdgeXpert platforms now support NVIDIA Nemotron 3.5 Lightning, a lightweight open model for agentic AI, and NeMo Switchyard, an open-source model routing …
CTO Advisor Keith Townsend says 95% of enterprise AI projects fail to return anything, and the failure occurs at the agentic AI reasoning plane, not at the model or data levels. Townsend, who moved a …
NVIDIA introduced Nemotron 3.5 Lightning, a customizable open 30B mixture-of-experts model for always-on agents, delivering up to 4x faster token generation and 30% faster time to completion compared …
Lucebox partnered with AMD to sell a $6,499 local AI inference computer combining a Radeon AI PRO R9700 GPU and Ryzen AI MAX+ 395 processor, achieving 51.1 tokens per second on a 284-billion-parameter…