09:35
2026-07-28
twitter.com
artificial-intelligence
Someone runs K3 on 80x 5090s, for 20 tok/s
A team has run the full Kimi K3 model, a 2.8-trillion-parameter mixture-of-experts (MoE) model, on 80 RTX 5090 GPUs achieving 20 tokens per second single-stream inference on day one without tuning, maโฆ