Someone runs K3 on 80x 5090s, for 20 tok/s A team has run the full Kimi K3 model, a 2.8-trillion-parameter mixture-of-experts (MoE) model, on 80 RTX 5090 GPUs achieving 20 tokens per second single-stream inference on day one without tuning, marking the first time frontier intelligence has been served using only GDDR7 gaming cards and plain ethernet with official MXFP4 weights. Kimi Moonshot released the model weights and technical report, stating the new architecture delivers 2.5x the intelligence per unit of compute and includes native visual understanding with a 1-million-token context window. we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s. 20 tok/s single stream, day one, untuned. Last week we took GLM-5.2 from 30 to 110 tok/s on this same fleet. This number will climb. A first for open weights: frontier intelligence served with zero HBM, the scarcest silicon in AI. Just GDDR7 gaming cards, plain ethernet, and the official MXFP4 weights, nothing requantized. The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it. @Kimi Moonshot https://x.com/Kimi Moonshot Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside