Kimi k3 now runs on one consumer GPU
Kimi's k3 model now runs on a single consumer GPU, achieving 113.83 tok/s on a 32 GB RTX 5090 after being shrunk to 28.8 GB from its original 48B size. The model works with coding agents like Claude C…
Kimi's k3 model now runs on a single consumer GPU, achieving 113.83 tok/s on a 32 GB RTX 5090 after being shrunk to 28.8 GB from its original 48B size. The model works with coding agents like Claude C…
A team has run the full Kimi K3 model, a 2.8-trillion-parameter mixture-of-experts (MoE) model, on 80 RTX 5090 GPUs achieving 20 tokens per second single-stream inference on day one without tuning, ma…
Kimi K3, the world's largest open-source model with 2.8 trillion parameters and a Mixture-of-Experts architecture requiring all 896 experts to be loaded into VRAM, needs a minimum of 6× H100 80GB GPUs…
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 using Unsloth's UD-IQ4_NL_XL quantisation achieved up to 140 tokens per second for generation and over 3,300 tok/s for prompt processing with a…
Black Forest Labs and mimic robotics have deployed FLUX-mimic, a video-action model that controls factory robots at Audi with a 101ms full-system reaction time, using a single NVIDIA RTX 5090 GPU. The…
A $1,799 Mac Mini M4 outperforms the $4,000 RTX 5090 for most local AI workloads in 2026, according to a comparison by an unnamed analyst. The Mac Mini's 48GB unified memory exceeds the RTX 5090's 32G…
KTransformers, a flexible LLM inference framework, enables deployment of 100B+ parameter models locally on a single RTX 5090 (32GB VRAM) using CPU/GPU heterogeneous computing without quantization. The…
Hugging Face has integrated Nunchaku 4-bit diffusion inference natively into Diffusers, enabling users to load quantized checkpoints with a simple from_pretrained() call and no local CUDA compilation.…
Poolside's Laguna S 2.1, a 118B MoE coding model, runs on a single RTX 5090 (32 GB) at ~19 tok/s decode and ~60 tok/s prefill using auto-fit layer placement. A developer achieved this by packing full …
A developer achieved ~28 tok/s single-stream decode of DeepSeek-V4-Flash (284B MoE) on a single RTX 5090 with dual Xeon Gold 6138 CPUs using a CPU-offloaded expert recipe. The setup uses ktransformers…
Nvidia at SIGGRAPH in Los Angeles detailed DLSS 5, its AI-powered graphics tool, promising to preserve 'artistic intent' after backlash at GTC earlier this year. The company demonstrated three model o…
Nvidia has not announced a successor to its RTX 5090 graphics card more than a year after its launch, and AI models ChatGPT, Perplexity, and Gemini suggest the company's focus on higher-margin AI and …
PrismML shipped Bonsai 27B, a 27B-parameter model based on Qwen3.6 27B, with a 1-bit variant packing to 3.9 GB that fits on an iPhone 17 Pro, clearing the memory gate that blocked prior builds of this…
Kalshi, the first federally regulated prediction market exchange in the US, has launched contracts allowing traders to speculate on the per-hour cost of running NVIDIA's H100, H200, B200, and RTX 5090…
VRAM capacity and memory bandwidth, not AI TOPS, are the decisive specs for local LLM inference in mid-2026, according to an analysis by Ji-ho Choi. NVIDIA's RTX 5090 leads with 32 GB VRAM and 1792 GB…
Two used RTX 3090s offer 48GB of VRAM for about half the cost of a single RTX 5090, enabling local 70B LLM inference at 14-21 tokens/sec, while the 5090's 32GB cannot run a quality 70B model. For mode…
AI builder Alex Finn runs a 24/7 local AI setup consisting of three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a self-built fleet dashboard. Finn spe…
The memory industry is emerging as a structural beneficiary of the AI era as autonomous AI agents and automation platforms drive demand for higher memory bandwidth, according to a Vault Track analysis…
A hyper-optimized, zero-dependency C/CUDA inference engine for the Qwen 3.6 35B model on RTX 5090 Blackwell GPUs achieves 13.4k tokens/sec prefill throughput at 2,048 context depth and 270+ tokens/sec…
NVIDIA's RTX 5090 offers 32GB of VRAM and 78% more memory bandwidth than the RTX 4090, but the extra performance is most noticeable for specific local LLM workloads such as running 32B models at Q4 wi…