Unsloth - Qwen3.8 - How to Run Locally
Unsloth released dynamic GGUF quantizations for Qwen3.8, enabling the 27B model to run locally on 17-19GB VRAM setups and the 2.4T parameter model to run in 397GB via 1-bit quantization. The Qwen3.8 f…
Unsloth released dynamic GGUF quantizations for Qwen3.8, enabling the 27B model to run locally on 17-19GB VRAM setups and the 2.4T parameter model to run in 397GB via 1-bit quantization. The Qwen3.8 f…
DeepSeek, Alibaba, and xAI all shipped new frontier models on August 13, 2026: DeepSeek V4 Pro 0813 (1.6T MoE, 49B active, 1M context, priced at $0.435/$0.87 per million tokens), Qwen3.8-2.4T-A95B (op…
Unsloth Desktop, a new local AI training tool from Unsloth, enables users to train models locally without a cloud backend, supporting NVIDIA, AMD, Intel, and Mac hardware. It claims a 70% reduction in…
Unsloth released Unsloth Desktop (Beta), a free, open-source app for running and training AI models locally on macOS, Windows, and Linux, supporting LLMs, diffusion image/video, MLX, GGUF, and audio m…
NVIDIA introduced Nemotron 3.5 Lightning, a customizable open 30B mixture-of-experts model for always-on agents, delivering up to 4x faster token generation and 30% faster time to completion compared …
NVIDIA's open-weights Nemotron 3.5 Lightning mixture-of-experts model, with 30 billion total parameters but only 3 billion active, is designed for agentic grunt work and can be run locally on consumer…
Qubitz, a local-first AI agent for GGUF models on llama.cpp, aims to make 7B–35B MCP/tool-capable LLMs more predictable and useful through a specialized harness and Agent Behavioral Contracts. It oper…
The Watershed moment in AI is not a single event but four distinct watersheds at different scales and price points, according to a new analysis. The first, the Frontier Watershed, occurred in November…
DeepSeek's V4 Flash 0731, a 284B-parameter Mixture-of-Experts model with 13B active parameters and a native 1M-token context under an MIT license, is cheaper to use via the API than to run locally, ac…
Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter sparse Mixture-of-Experts model, on August 3, claiming it outperforms GPT-5.6 Sol on key coding benchmarks and promising full open weights next w…
Independent researcher Alpamys Makazhan released Soup, a Show HN project that fine-tunes a full Llama-3.1-8B model in NF4 quantization with a 3.32 GB VRAM peak at 119.6 tokens/sec on a 4 GB RTX 3050 L…
Unsloth, an open-source AI startup, released Unsloth Studio (Beta), a platform that lets users run and train text, audio, embedding, and vision models locally on Windows, Linux, and macOS, with suppor…
DeepSeek released V4 Flash 0731, an iterative update of its Mixture-of-Experts model with 284B total parameters and 13B active per token, achieving an Artificial Analysis Intelligence Index of 50, up …
Z.ai's GLM-5.2 scored 62.1 on SWE-bench Pro, beating GPT-5.5's 58.6, while costing $1.40 per million input tokens and $4.40 per million output tokens, making it 3.6x cheaper on input and 5.7x cheaper …
DeepSeek released DeepSeek-V4-Flash-0731 on Hugging Face and moved its V4-Flash API to public beta on July 31, 2026, claiming major gains in agentic and coding benchmarks from re-post-training rather …
The accuracy gap between the best open-weight models and GPT-4o has shrunk to under 3% on structured data tasks, according to a developer's hands-on analysis. A fine-tuned Qwen2.5 72B model achieved 9…
A first-principles analysis of LoRA fine-tuning memory usage reveals that 87.3% of VRAM scaling with sequence length for Llama 3.1 8B is consumed by the cross-entropy loss head tensor, not the model, …
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 using Unsloth's UD-IQ4_NL_XL quantisation achieved up to 140 tokens per second for generation and over 3,300 tok/s for prompt processing with a…
Unsloth achieves up to 7.3x training speedup over standard Transformers for MoE models like gpt-oss-20B on an NVIDIA B200, according to published benchmarks, while Axolotl delivers up to 1.45x speedup…
AMD's Ryzen AI Developer Center, bundled with Ryzen AI Max Plus 395 machines, eliminates manual ROCm and driver setup for local AI workloads, offering guided playbooks for ComfyUI, LM Studio, and Unsl…