cd /news/artificial-intelligence/ran-moonshot-s-2-8t-parameter-kimi-k… · home topics artificial-intelligence article
[ARTICLE · art-78162] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Ran Moonshot's 2.8T-parameter Kimi K3 on a GPU-less mini-PC

Moonshot AI published Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model, on July 27, 2026, and rabbit ran it the next day on a GPU-less mini-PC using a Rust engine that streams routed experts from disk on demand, achieving correct output for a sample prompt in 2698.1 seconds for 40 tokens with no performance tuning yet.

read7 min views5 publishedJul 29, 2026
Ran Moonshot's 2.8T-parameter Kimi K3 on a GPU-less mini-PC
Image: source

Small hops, immense models. A Rust engine that runs frontier open-weight MoE models on a single machine. The dense part stays resident in RAM; the routed experts stream from disk on demand.

Moonshot AI published Kimi K3 on 2026-07-27, a 2.8-trillion-parameter Mixture-of-Experts model and one of the largest open-weight releases so far. rabbit runs it the next day, straight off Moonshot's published checkpoint. No conversion step.

$ rabbit --model /mnt/data/kimi-k3 --prompt "What is the capital of France?" --max-tokens 40
 model (dbits=4, ebits=4)...
model loaded in 610.0s (93 layers, 896 experts/layer)
prefill (7 tokens)...
  prefill done in 412.8s
  ...
...response["answer"] == "Paris"...

40 tokens in 2698.1s

Correctness validated three ways before that run. Bit-exact against a random-weight instance of Moonshot's own PyTorch reference (teacher-forcing, every position, tests/teacher_forcing_k3.rs

). A structural smoke test against the real 1.56TB checkpoint (examples/k3_smoke.rs

). And the real prompt above, answered correctly. No performance work has landed for K3 yet; the numbers above are correctness-first, not tuned. rabbit's other architecture, GLM-5.2, started from a similarly slow floor and is now 3.5× faster across eight measured versions. See Honest numbers and PERFORMANCE.md

.

K3 brings Kimi Delta Attention plus a Gated MLA hybrid, Stable LatentMoE (routed experts compute in a narrower latent width, not the full hidden size), Attention Residuals (a transient block-pooling mechanism across layers), and native OCP MXFP4 quantization for its 896 experts/layer. rabbit reads that MXFP4 straight off disk, byte for byte, no requantization.

A large MoE model activates a small fraction of its parameters per token, and only the routed experts change token to token. So, across every architecture rabbit runs:

  • the dense part(attention, shared experts, embeddings) stays** resident in RAM**, quantized; - the routed experts, thousands of them, tens of MB each, live** on diskand are streamed on demand**, through a per-layer LRU cache, a persistent learned pin for the hottest ones, and the OS page cache as a free extra tier.
Kimi K3 Kimi Linear 48B GLM-5.2
total / active params 2.8T / not yet characterized 48B / ~3B 744B / ~40B
attention KDA + Gated MLA, extra output gates Kimi Delta Attention + Gated MLA MLA + DSA sparse indexer
MoE routing Stable LatentMoE (narrower latent width), shared experts grouped routing, shared experts noaux_tc sigmoid, shared expert
native quantization read OCP MXFP4 (routed experts) BF16 FP8 (E4M3, block-scale)
checkpoint Moonshot's real release Moonshot's real release pre-converted by colibrì's tooling
status correctness-validated, perf work pending tuned (--session , real chat)
fully tuned, 8 versions of perf work

One Model

/KvState

/ExpertCaches

family-dispatch enum in src/model.rs

routes to the right architecture from config.json

's model_type

. --chat

/--serve

/--prompt

/--session

all work the same way across the three.

Faithful forward pass for all three architectures. Validated token-exact against a synthetic oracle built from each model family's own real reference code, plus real-checkpoint validation for each.MLA attention withweight absorption for decode (no per-token k/v reconstruction) and dense reconstruction for prefill, parallelized withrayon

across attention heads.DSA sparse attention(GLM-5.2's lightning indexer) and** Kimi Delta Attention**(KDA's chunked recurrence, short convolutions, per-channel decay gate). Real math, not approximated.** int4/int8/int2 quantization**, native** FP8**(E4M3, block-scale) and** OCP MXFP4**checkpoint , grouped-scale int4, and apre-quantized fast path..qs

AVX2 + AVX-512/VNNI kernels, runtime-selected, plusrayon

parallelization across CPU cores for every matmul and the absorbed-attention decode path., with a sequential-io_uring

-batched expert streamingpread

fallback. K3's MXFP4 experts use the fallback today; batching that path is open work (seeROADMAP.md

).Persistent expert usage cache(.rabbit_usage

). Learns which experts your usage routes to and pins them, lazily, once a candidate is actually loaded through normal use.KV-cache persistence(--session

). Conversations reopen warm across restarts.A standalone checkpoint converter(bin/convert.rs

). Architecture-agnostic tensor classification, per-bucket bit-depth control, a--report

quality pass. No dependency on colibrì's own tooling except for the pre-converted GLM-5.2 checkpoint above.OpenAI-compatible HTTP server(--serve

). Streaming and non-streaming/v1/chat/completions

,/v1/models

, plus/profile

, a rolling per-turn phase-timing window.Multi-turn chat(--chat

) with each model's own real chat template.

Not yet built: io_uring

-batched MXFP4 expert , a SIMD tier for the MXFP4 matmul kernel (scalar only today), live expert re-pinning, GPU/CUDA, MTP speculative decoding, ARM NEON, grammar-constrained decoding, and a web UI (/profile

is a JSON endpoint, no page serves it yet; see DASHBOARD_BRIEF.md

).

metric value
model load 610.0s (93 layers, 896 experts/layer)
prefill (7 tokens) 412.8s
decode, steady state 50-70s/token
decode I/O share 30-40% disk I/O, 60-70% compute
expert-cache hit rate (40-token run, --expert-cache 64 )
grows to ~52% cumulative, still ~43% miss rate late in the run

This is the correctness-first floor, not a tuned number. Same starting point GLM-5.2 was at before its own eight versions of rayon

/SIMD/io_uring

work. The single biggest known lever: the MXFP4 matmul kernel is scalar only, and compute, not disk, is most of each token's time here, the opposite of GLM-5.2's I/O-bound steady state below.

metric value
checkpoint 378 GB (jlnsrk/GLM-5.2-colibri-int4 )
rayon matmul parallelization
128.9s → 36.3s for 5 tokens (3.5×), bit-exact output
rayon absorbed-attention parallelization
224.3s → 158.4s for 70 decode tokens (~29% faster)
decode I/O share, steady state (warm cache) 30-35% disk I/O, 65-70% compute
prefill I/O share (cold cache) ~75% disk I/O
expert-cache hit rate, steady-state decode 70-77% (miss floor 23-30%)
usage-cache auto-pin 150 experts (2/layer × 75 MoE layers): prefill hits 0 → 136
decode speed, current (v0.22.0) 1.02 words/sec, up from 0.29 across eight measured versions

All measured against the real checkpoints, not estimated. See PERFORMANCE.md

for the full chronological log, including techniques that were tried and reverted.

cargo build --release
cargo test
./target/release/rabbit --model /mnt/data/kimi-k3 --prompt "What is the capital of France?"
./target/release/rabbit --model /mnt/data/kimi-k3 --chat --session ~/.rabbit_session
./target/release/rabbit --model /mnt/data/kimi-k3 --serve --port 8000

--model

auto-detects the architecture from the checkpoint's config.json

. Kimi K3 and Kimi Linear read Moonshot's own published safetensors

shards directly: download moonshotai/Kimi-K3 and point

--model

at it, no conversion step. GLM-5.2 needs a directory in colibrì's converted layout; a pre-converted checkpoint such as works directly. See

jlnsrk/GLM-5.2-colibri-int4

--help

for the full flag list (--max-tokens

, --temperature

, --nucleus

, --expert-cache

, --dbits

, --ebits

, --shard-dirs

, --no-usage-cache

, ...).

src/
├── kimi_k3/                                  Kimi K3: SituAndMul, LatentMoE, Attention Residuals, MXFP4
├── kimi_linear/                              Kimi Linear 48B: KDA, short convs, tokenizer, chat template
├── glm52/                                    GLM-5.2: MLA+DSA attention, MoE router, checkpoint converter
├── model.rs                                  family-dispatch enum: Model/KvState/ExpertCaches/Tokenizer
├── safetensors.rs, quant.rs, kernels.rs      shard index, quantization, scalar/AVX2/AVX-512/MXFP4 kernels
├── expert_cache.rs, usage_cache.rs           LRU expert streaming + persistent usage learning
├── generate.rs, kv_session.rs                shared generation loop + KV-cache persistence
├── chat.rs, server.rs, main.rs               chat templates, HTTP server, CLI entrypoint
tests/oracle/     per-architecture oracle generators (real reference code, vendored) + fixtures
tools/            real-tokenizer validation fixtures (dev-only, not a runtime dependency)
benches/          criterion benchmarks (kernels, expert )

colibrì is C, hand-written, effectively zero-dependency, and GLM-5.2-only. rabbit ports the same algorithms to Rust with a short list of well-justified dependencies instead of a zero-dep stance, adds its own performance work (the rayon

parallelization above, KV-session and expert-usage persistence), and generalizes the whole engine into a family-dispatch design that now runs two more architectures colibrì doesn't. Every architecture is validated the same way colibrì validates itself: token-exact teacher-forcing against a tiny synthetic model built from that architecture's own real reference code.

Part of the ferrumox AI lab, alongside fox (a production local-LLM server wrapping llama.cpp). rabbit is the opposite kind of project: a research engine for models that don't fit in memory even offloaded, built by hand instead of wrapping an existing runtime.

The name is a nod to RabbitLLM, an earlier, unrelated project of mine (a fork of AirLLM that streams full model layers through limited GPU VRAM). Same interest in running large models on constrained hardware, different problem, a completely different technique. Nothing in this codebase is derived from that one.

Pre-1.0, 0.MINOR.PATCH

, tracked via git tags and release/vX.Y.Z

branches rather than Cargo.toml

's version field. MINOR

bumps at the end of each development phase, PATCH

for fixes/polish within one. rabbit-plan.md

has the full phase-by-phase history through GLM-5.2 and Kimi Linear's bring-up.

TBD.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @moonshot ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ran-moonshot-s-2-8t-…] indexed:0 read:7min 2026-07-29 ·