cd/entity/RTX 5090· home entities RTX 5090
grep -l @rtx 5090 /news/*.json | wc -l → 72

RTX 5090

mentions 72 type Person page 2/4 feed RSS

// recent coverage 72 mentions

20:02
2026-07-29
twitter.com
large-language-models

Kimi k3 now runs on one consumer GPU

Kimi's k3 model now runs on a single consumer GPU, achieving 113.83 tok/s on a 32 GB RTX 5090 after being shrunk to 28.8 GB from its original 48B size. The model works with coding agents like Claude C…

09:35
2026-07-28
twitter.com
artificial-intelligence

Someone runs K3 on 80x 5090s, for 20 tok/s

A team has run the full Kimi K3 model, a 2.8-trillion-parameter mixture-of-experts (MoE) model, on 80 RTX 5090 GPUs achieving 20 tokens per second single-stream inference on day one without tuning, ma…

15:12
2026-07-27
openmodelmap.com
large-language-models

Kimi K3 Hardware Requirements

Kimi K3, the world's largest open-source model with 2.8 trillion parameters and a Mixture-of-Experts architecture requiring all 896 experts to be loaded into VRAM, needs a minimum of 6× H100 80GB GPUs…

22:33
2026-07-24
gilesthomas.com
large-language-models

Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090

Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 using Unsloth's UD-IQ4_NL_XL quantisation achieved up to 140 tokens per second for generation and over 3,300 tok/s for prompt processing with a…

16:08
2026-07-24
sourcefeed.dev
artificial-intelligence

FLUX 3's Real Headline Is the Robot, Not the Video

Black Forest Labs and mimic robotics have deployed FLUX-mimic, a video-action model that controls factory robots at Audi with a 101ms full-system reaction time, using a single NVIDIA RTX 5090 GPU. The…

00:01
2026-07-24
pub.towardsai.net
artificial-intelligence

Mac Mini M4 vs RTX 5090 vs Cloud GPUs for Local AI in 2026

A $1,799 Mac Mini M4 outperforms the $4,000 RTX 5090 for most local AI workloads in 2026, according to a comparison by an unnamed analyst. The Mac Mini's 48GB unified memory exceeds the RTX 5090's 32G…

09:03
2026-07-23
ktransformers.net
artificial-intelligence

KTransformers – Flexible LLM Inference Framework

KTransformers, a flexible LLM inference framework, enables deployment of 100B+ parameter models locally on a single RTX 5090 (32GB VRAM) using CPU/GPU heterogeneous computing without quantization. The…

00:00
2026-07-23
huggingface.co
artificial-intelligence

Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

Hugging Face has integrated Nunchaku 4-bit diffusion inference natively into Diffusers, enabling users to load quantized checkpoints with a simple from_pretrained() call and no local CUDA compilation.…

23:05
2026-07-22
gist.github.com
large-language-models

Run Poolside Laguna S 2.1 (118B MoE) on a single RTX 5090

Poolside's Laguna S 2.1, a 118B MoE coding model, runs on a single RTX 5090 (32 GB) at ~19 tok/s decode and ~60 tok/s prefill using auto-fit layer placement. A developer achieved this by packing full …

19:04
2026-07-14
sourcefeed.dev
artificial-intelligence

Bonsai 27B Puts Real Agents on Phones

PrismML shipped Bonsai 27B, a 27B-parameter model based on Qwen3.6 27B, with a 1-bit variant packing to 3.9 GB that fits on an iPhone 17 Pro, clearing the memory gate that blocked prior builds of this…

14:16
2026-07-14
cryptobriefing.com
artificial-intelligence

Kalshi builds prediction markets for GPU computing power prices

Kalshi, the first federally regulated prediction market exchange in the US, has launched contracts allowing traders to speculate on the per-hour cost of running NVIDIA's H100, H200, B200, and RTX 5090…

04:02
2026-07-14
sourcefeed.dev
artificial-intelligence

VRAM Beats TOPS for 2026 Local AI GPUs

VRAM capacity and memory bandwidth, not AI TOPS, are the decisive specs for local LLM inference in mid-2026, according to an analysis by Ji-ho Choi. NVIDIA's RTX 5090 leads with 32 GB VRAM and 1792 GB…

06:22
2026-07-13
calcrecipe.com
artificial-intelligence

The Winners of the AI Era

The memory industry is emerging as a structural beneficiary of the AI era as autonomous AI agents and automation platforms drive demand for higher memory bandwidth, according to a Vault Track analysis…

← prev page 2 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics