cd /news/large-language-models/show-hn-run-full-kimi-k3-with-29-gb-โ€ฆ ยท home โ€บ topics โ€บ large-language-models โ€บ article
[ARTICLE ยท art-79784] src=github.com โ†— pub= topic=large-language-models verified=true sentiment=โ†‘ positive

Show HN: Run Full Kimi K3 with 29 GB of RAM

A developer has demonstrated running the full 2.78-trillion-parameter Kimi K3 model on a consumer 64 GB MacBook Pro using a custom inference engine called WASTE, achieving 0.32โ€“0.34 tokens per second with a minimum RAM requirement of 29.05 GB. The engine streams experts from disk and uses RAM as a bounded cache, marking the first known instance of a trillion-scale model streaming from NVMe on consumer hardware without distillation or pruning.

read26 min views1 publishedJul 30, 2026
Show HN: Run Full Kimi K3 with 29 GB of RAM
Image: source

Kimi K3 โ€” 2.78 trillion parameters โ€” running on a consumer laptop.

$ waste run ~/models/k3.waste 'What is the capital of Italy?'
waste: no --budget, using 46.24 GB of 64.00 GB (expert cache 17.56 GB)
The capital of Italy is **Rome**.
[16 tokens, 49.31 s, 0.32 tok/s | experts 3357 hit / 20195 miss = 14%]

WASTE is an embeddable inference engine written in C, with no third-party runtime dependencies. It keeps the model trunk in memory, streams selected experts directly from disk, and uses the remaining RAM as a bounded expert cache.

Its current proof point is the complete open-weights Kimi K3 model: 2.78 trillion parameters, converted into a 982 GiB container and running on a 64 GB MacBook Pro at 0.32โ€“0.34 tokens per second. This is not a distilled, pruned, or reduced variant.

Model Container Minimum RAM Tested speed
Kimi K3 2.78T
982 GiB 29.05 GiB 0.32โ€“0.34 tok/s
Kimi-Linear 48B
19 GiB 1.86 GiB 8.92 tok/s

WASTE was written for that one model and that one constraint: K3 does not fit in the RAM of current mainstream consumer systems. It is 1.42 TB as published and 982 GB after conversion. But a mixture of experts activates about 4% of itself per token, so almost all of that weight is idle at any instant โ€” and idle weight does not need to be in memory, it needs to be reachable in time. WASTE keeps it on disk in a layout where one expert costs exactly one read, streams what each token actually needs, and spends every remaining byte of RAM on the part that repeats.

The engine is correct: every layer is validated against a PyTorch reference, the final logits agree to 3.6e-06, and the vision tower matches its own oracle to 2.3e-06. It is also slow โ€” a third of a token per second, fifty seconds for the sentence above.

Both of those matter, and the second one should not be read as a disclaimer. We are not aware of another published demonstration of a model this size streaming from disk on a consumer machine: we found none for trillion-scale NVMe streaming, and the best-documented 671B-class recipes assume a server with a terabyte of DDR5. That is a report of what our search turned up rather than a survey โ€” this repository carries no bibliography and no comparison table, so read it as an invitation to send a counter-example, not as a result. The interesting part is not the speed, it is that the whole thing is in the reachable range on a single consumer machine โ€” and that from here the question is engineering rather than feasibility. Half of a decode step is already disk I/O running near the drive's measured ceiling, so the levers are known: read fewer bytes per token, and keep more of them in RAM.

What that opens up, concretely: a frontier-scale model that answers with no network, no per-token invoice, and nothing leaving the machine โ€” which is the difference between "you may not send that data to an API" and "run it here". The format and the engine are not K3-specific in any deep way; K3 is simply the hardest case that exists today, and a model that streams at 2.78T streams comfortably at 48B.

Every number in this document was measured on the commit it is published with, and the ones that were wrong are recorded as wrong in docs/LEARNED.md rather than quietly corrected.

Every token answered by a cloud service is paid for twice: once on the invoice, and once in the electricity of a datacenter running a model that would fit โ€” barely, awkwardly, but genuinely โ€” on hardware already sitting on a desk. WASTE means to be the first concrete step toward ending that waste of tokens. The acronym came second.

disk, for the model | 982 GB for the converted container โ€” plan a terabyte | | disk, to convert it | another 1.42 TB of staging for the published shards, freed afterwards | RAM | 29.05 GB minimum to open K3 at 4K context; 64 GB for the numbers here | | storage speed | the container must be on internal NVMe โ€” see below | | build | a C11 compiler and make . No BLAS, no CUDA, no Python at run time |

Sizes here are powers of two, the way df

and the engine both report them: the container is 982 GiB, which a disk vendor would call 1.05 TB.

The RAM floor is what the engine refuses to start below, and it is almost entirely the 27.28 GB resident trunk. Useful throughput starts higher: on a 64 GB machine the engine gives itself a 46 GB budget, of which 17.56 GB is expert cache, and that is the top of the measured curve. A 32 GB machine can technically open the model and will page badly; treat 64 GB as the real requirement.

Storage speed is not a detail. A token reads 17 GB of experts. On the internal SSD that is 12.78 GB/s and the model streams; over a USB enclosure it is 0.94 GB/s and the same token takes thirteen seconds. Convert onto internal NVMe, and use the external disk for the download only.

If a terabyte is not available, the same engine and the same format run Kimi-Linear-48B-A3B-Instruct

from a 19 GB container with a 1.86 GB floor, at 8.92 tok/s. That is the good path for trying WASTE out before committing a disk to K3.

Self-contained. Onelibwaste.a

, onewaste

binary, nothing at run time beyond libc and pthreads.Zero dependencies. No BLAS, no ONNX, no Python in the inference path, nothing to install. The Python undertools/

converts models and validates the engine; it never runs alongside it.Fully embeddable. Twenty-six public functions insrc/waste.h: open a model under a RAM ceiling, generate, save the session, close. The CLI is a client of that API and touches nothing private โ€” if the CLI can do it, so can an embedding host.

waste_cfg cfg;
waste_cfg_init(&cfg);
cfg.ram_budget_bytes = 46ULL << 30;  /* a hard ceiling, not a hint;
                                        0 sizes it to this machine */

waste_ctx *ctx;
if (waste_open("/path/to/k3.waste", &cfg, &ctx) != WASTE_OK) return 1;
waste_generate(ctx, ids, n, &params, on_token, user);
waste_close(ctx);

The path is the container directory the converter wrote โ€” no ~

expansion here, that is the shell's job.

A model is converted once into a .waste

container: a JSON manifest, a resident trunk, and one expert bank per layer. Each expert record is 4 KiB-aligned with its gate, up and down matrices adjacent, so routing to an expert costs exactly one pread โ€” not three, not a seek per matrix. The arithmetic was never the bottleneck.

Reads bypass the page cache (F_NOCACHE

on macOS, O_DIRECT

on Linux, FILE_FLAG_NO_BUFFERING

on Windows). That is deliberate: with a container smaller than RAM the kernel would cache everything, and the hit rates measured that way are a fiction that does not survive contact with a 982 GB model.

Every record's header is checked on the way in โ€” right magic, the expert the index asked for, offsets that fit โ€” so a bank that has been truncated or spliced stops the generation and names the record instead of answering from the wrong bytes. That costs nothing measurable. The record also carries a crc32

over its payload, and checking that is --verify

, off by default: it is a pass over every record on every cache miss, about 5% on Kimi-Linear and 1% on K3. Worth it for a container you copied or downloaded and have not read since; not worth it on every token of one you converted yourself. See docs/FORMAT.md.

Experts are stored as residual vector quantization โ€” three stages of 256-entry codebooks over 8-dimensional vectors, 3.00 bits per weight โ€” and the matrix is never materialized. For each token the engine builds a table of partial dot products, one per codebook entry per vector position, after which every expert row is three table reads and two adds.

The trunk stays at 4 and 8 bits. The model was trained with quantization-aware training on the experts only, so it has no trained tolerance for a squeezed trunk: a 3-bit trunk was built and measured, the cache prediction held, the throughput did not, and the output collapsed.

The most predictive number in this project. K3 touches 16 experts in each of 92 layers per token: 17.0 GB. Below that, an expert cached for one token is evicted before the next token asks for it, and the hit rate is not low โ€” it is zero. Above it the curve bends sharply.

budget expert cache hit rate decode
32 GB 3.32 GB 0% 0.31 tok/s
46 GB 17.32 GB 13% 0.32 tok/s
52 GB 23.32 GB 27% 0.11โ€“0.14 tok/s
58 GB 29.32 GB 37% 0.04 tok/s

Measured in that order, on an otherwise idle machine. Order matters: re-run after the 52 and 58 GB rows have driven the machine into paging, 46 GB gives 0.22โ€“0.25 rather than 0.32 โ€” while reporting hit and miss counts identical to the digit. The engine is deterministic; the machine is not, and it does not fully recover between runs. Sweep upward.

Everything in the memory design exists to get above that line, which is why the engine works to free RAM rather than to save it.

And there is a ceiling on the other side, closer than it looks. Read that table twice: the hit rate climbs all the way down. At 58 GB on a 64 GB machine the cache serves 37% of experts from RAM and the engine is eight times slower than at 46 GB, where it serves 13%. The engine is inside its budget; the machine is not, so the OS pages out the expert cache, and a "hit" becomes a page fault instead of the disk read the engine was managing.

So the usable window is narrow. It opens at ~46 GB, where the cache finally clears one token's working set, and it has already closed by 52 โ€” on an otherwise idle machine, with 49 GB free before the run. It is also sharp enough to move under a change that looks unrelated: taking 1.11 GB of embedding table off the resident set fed straight into the cache at a fixed budget, and that was enough to push 58 GB from 0.32 tok/s to 0.04.

So the default does not fill the machine. Expert cache is only worth anything in whole multiples of that working set, and the remainder above a multiple buys a few points of hit rate while pushing the machine towards paging. When it picks a budget for itself the engine steps down a whole working set at a time and takes the largest that fits under seven eighths of RAM: K3 asks for floor + 3ร— โ€” 80.63 GB โ€” and gets floor + 1ร— on this laptop, a 46 GB budget and a 17.56 GB cache. That is the top of the curve above, reached with no flag. A 128 GB machine still gets the full 3ร—.

An earlier version took every byte up to the cap instead, which put a 27 GB cache on this machine โ€” between two budgets measured at 0.11 and 0.04 tok/s. The real lesson is that a cache you do not control is not a cache, and the corollary is that an engine should stop asking for memory before the OS starts taking it back.

K3's attention is a 3:1 hybrid: Kimi Delta Attention, which carries a fixed-size recurrent state instead of a growing KV cache, and gated multi-head latent attention. The MLA layers cache the 512-wide latent rather than expanded per-head keys and values, with kv_b_proj

absorbed into the query and the output:

q_nope ยท (W_kb c)    ==  (W_kbแต€ q_nope) ยท c
ฮฃ_s a_s (W_vb c_s)   ==  W_vb (ฮฃ_s a_s c_s)

Identical logits to 1.2e-05, and 53ร— less cache: 11.25 GB becomes 0.21 GB at 4K context. It is also what makes long context possible at all โ€” the expanded layout wants 360 GB at 128K tokens, the latent one 7.2.

MacBook Pro M5 Pro, 64 GB, container on the internal SSD. Every figure was measured on the commit it is published with.

| minimum RAM | 29.05 GB at 4K context | | 30.54 GB at 32K, 35.63 GB at 128K, 83.21 GB at 1M | | | resident trunk | 27.28 GB | | read per token | 17.0 GB at ~9.9 GB/s, near the SSD's measured ceiling | | model load | 20 s | | prefill | 0.47 tok/s chunked, 0.29 sequential | | decode | 0.32โ€“0.34 tok/s at the default budget, the best this machine gives | | vision tower | 15.7 s for a 1024-patch image, 27 layers | | image in a prompt | 256 positions for 896x896, 2.8 s each โ€” as text |

The floor is almost entirely the resident trunk. Useful throughput starts above ~46 GB, where the expert cache finally clears one token's working set, and is gone again by 52, where the machine starts paging. Below the first line extra RAM buys nothing; above the second it costs, badly. The window is one budget wide on this machine.

The tower is not what an image costs. Encoding 1024 patches takes 15.7 s; the 256 positions it produces then go through the 92 MoE layers like any other token, which is the other 731 s. An image is priced as text of the same length, so the patch budget in vision.json

is a real dial: halving the grid halves the prompt.

| minimum RAM | 1.86 GB | | decode | 8.92 tok/s at an 8 GB budget, 78% cache hit |

The same engine and the same format, on a model that fits comfortably. This is what WASTE looks like when it is not fighting.

Decode on K3, 17.32 GB of cache and still cold โ€” 6.7% hit over ten steps, which is the state a fresh prompt starts in:

share
MoE, all of it 82.5%
of which expert I/O 53.5%
of which expert matmul 20.0%
KDA layers 14.5%
MLA layers 2.8%
lm_head 0.2%

Reproduce with WASTE_PROFILE=1 WASTE_CACHE_MB=17735 ./test_forward MODEL 1008,10484,318,15383,387 out.bin 5

. The I/O share falls as the cache warms, so a long session sits lower than this; the ranking does not change.

The I/O already runs near the hardware limit โ€” 17.0 GB per token at ~9.9 GB/s against the SSD's measured 12.78 โ€” so it only gets cheaper by happening less often, which means cache, which means RAM. That is the whole optimization story so far, and the reason the next steps are about memory rather than arithmetic.

git clone https://github.com/sqliteai/waste && cd waste
make                          # libwaste.a, waste, libwastevq
make check                    # 23 pass, 11 skip on a fresh clone

No configure step and no dependency resolution. make check

needs no model: it builds a small synthetic container and runs the engine against it. The eleven skips are the checks that need something a clone does not carry โ€” the PyTorch oracle, the round-trip against the source shards, anything driving the CLI with text, since the synthetic container carries no tokenizer, and the K3 checks, which want the container and the release on disk. With both containers present the suite is 36 checks.

Conversion is the one step that needs Python, and it happens once. The source is moonshotai/Kimi-K3 exactly as published โ€” 96 safetensors shards, 1.42 TB, nothing patched:

tools/fetch_weights.sh --dest /Volumes/staging/k3 --dry-run

tools/fetch_weights.sh --dest /Volumes/staging/k3

uv run --with torch --with safetensors python tools/convert.py \
    --src /Volumes/staging/k3 \
    --out ~/models/k3.waste --jobs 3

That produces the 982 GB container every number above was measured on. It takes about 4.7 hours with three processes on the M5 Pro (23.7 with the pure-torch encoder โ€” see docs/K3.md), and wants ~1.0 TB free on the target volume. The converter is resumable too: a layer whose bank is already written is skipped, so an interrupted run costs only the layer it was in the middle of.

The download is the part that goes wrong. A 1.42 TB pull over hours will hit dropped connections, CDN 5xx and at least one interrupted run, so every shard resumes mid-file rather than restarting, retries with exponential backoff and jitter, and counts as done only when its size matches Content-Length โ€” recorded in a state file, so a re-run skips finished shards without even a HEAD request. --check

re-verifies everything on disk against the remote and downloads nothing (96 shards in 34 s). --repo

points it at another model, HF_TOKEN

at a gated one. macOS and Linux.

Give --dest

a staging disk rather than the volume that will hold the container. The shards are read once, by the converter; the container is read continuously, at every token. On this machine the external enclosure measures 0.94 GB/s against the internal NVMe's 12.78 โ€” see docs/GATES.md, Gate H โ€” which is the difference between a model that streams and one that stalls.

tools/pipeline.sh

chains the whole thing unattended โ€” download, convert, round-trip the container against the source weights, generate, then diff the logits against the PyTorch oracle โ€” and leaves a report next to the container. The same converter handles the other member of the family, Kimi-Linear-48B-A3B-Instruct

, into the 19 GB container of the second benchmark; --src

is the only thing that changes.

Pre-converted containers are on their way to huggingface.co/sqliteai, at which point this whole section becomes a download and the Python is only needed for models we have not published.

The container is the directory the converter wrote, so give it that path โ€” ~/models/k3.waste

throughout this README:

waste run   ~/models/k3.waste "The capital of France is" -n 32
waste chat  ~/models/k3.waste                     # multi-turn, state kept
waste eval  ~/models/k3.waste "2 + 2 =" --top-k 5 # next-token distribution
waste plan  ~/models/k3.waste --budget 46G        # what fits, what does not
echo "prompt" | waste run ~/models/k3.waste       # stdin works too

-n

is a cap, not a requirement: without it generation stops at the container's end-of-sequence token or at 128 tokens, whichever comes first. The examples pass it because 128 tokens of K3 is six minutes.

--budget

is optional, and leaving it out is the right default rather than a fallback: the engine takes the container's recommendation, steps it down a whole token working set at a time until it fits under seven eighths of physical RAM, and never goes below the floor โ€” a budget you set explicitly under the floor is refused rather than swapped into. It then says on stderr what it landed on, so the same command on two machines is not silently two different runs:

waste: no --budget, using 46.24 GB of 64.00 GB (expert cache 17.56 GB)

--verify

checks each expert record's crc32

as it comes off the disk. It is off by default, and that is a throughput decision rather than a claim that containers do not rot: it is a pass over every record on every cache miss, about 5% on Kimi-Linear and about 1% on K3, where the read dominates. Turn it on once for a container you copied, downloaded, or left on a disk you do not trust, and for anything whose wrong answers would be believed; leave it off for one you converted yourself and have been reading since. WASTE_VERIFY=1

in the environment does the same thing, and the server takes --verify

as well. Any of them turns it on; none of them turns it off.

What is checked either way: a short read, and a record header that does not describe the expert the bank index asked for. Those are O(1), they cost nothing measurable, and they are what keeps a damaged offset out of the arithmetic โ€” --verify

only adds the pass over the payload.

waste --help

lists all nine commands. --json

makes eval

, tokenize

, plan

, info

and bench

machine-readable.

serve/

is an OpenAI-compatible HTTP server โ€” the second client of the public API, alongside the CLI, reaching the same engine through ctypes rather than keeping a copy of the model code in Python:

make libwaste.dylib                     # or libwaste.so on Linux
python3 -m serve ~/models/k3.waste --port 8000
curl localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"k3","messages":[{"role":"user","content":"Why is the sky blue?"}]}'

/v1/chat/completions

(streaming and not), /v1/completions

, /v1/models

, /health

. It carries the whole of K3's prompt format, not the four-string subset a container's chat.json

can hold: tool definitions and tool results, typed call arguments, JSON response schemas, tool_choice, the think channel and thinking_effort, and images โ€” plus the parser that reads the reply back into reasoning, answer and

tool_calls

. Stdlib only.The prompt renderer is a port of encoding_k3.py

from the release, and the test suite checks it against that file segment for segment on a corpus of 38 conversations whenever the weights directory is on disk. docs/SERVE.md is the reference.

K3 is multimodal โ€” a 401M ViT, 27 layers, patch 14 โ€” and so is the engine. --image

attaches a picture; repeat it for several:

$ waste run ~/models/k3.waste 'What is in this picture?' --image landscape.png
[landscape.png: 192 image tokens]
The picture shows a simple, stylized landscape with:

- A **blue sky** with a gradient from darker blue at the top to lighter blue near the horizon.
- A **yellow sun** in the upper right.
- A **gray hill or mountain** in the middle distance.
- A **green field** covering the lower part of the image.
[78 tokens, 234.25 s, 0.33 tok/s | experts 15314 hit / 99502 miss = 13%]

That is a 448ร—336 image, and every element of the description is in it โ€” including the sky gradient, which is the kind of detail that separates a tower that works from one that merely runs. The picture was generated by a twenty-line script rather than photographed, so the answer can be checked against what was drawn instead of against an impression.

PNG, JPEG, GIF, BMP, TGA and PSD, decoded by the one vendored header in third_party/

. It works on run

, chat

and eval

; inside a chat, /image FILE

attaches a picture to the next message, and it is spliced once โ€” the positions are in the attention state afterwards, so later turns discuss the same photograph without re-encoding it. The 27-layer ViT is loaded only when an image is present, because its 434 MB otherwise come straight out of the expert cache.

An image is not one token. The tower turns a 14-pixel patch grid into one embedding per merged 2ร—2 patch, and each occupies a position in the sequence โ€” the 448ร—336 above is 192 of them, an 896ร—896 photo at the default budget is 256. That is worth knowing before wondering where a context window went, and it is most of what an image costs: the 234 s in the transcript is the 78 generated tokens alone, and the picture is paid for before that, in prefill. An image is priced as text of the same length. The tower is the cheap part โ€” 15.7 s for a full 1024-patch image โ€” and its output then walks through 92 MoE layers like any other token. Halving max_patches

in vision.json

halves the bill.

Through the library it is three calls, because a host needs to size the prompt before committing to it:

size_t rows;
waste_image_add(ctx, "photo.png", &rows);          /* encode and queue   */
waste_image_expand(ctx, raw, n, ids, cap, &n_ids); /* placeholder -> N   */
waste_generate(ctx, ids, n_ids, &params, cb, u);   /* consumes the queue */

The tower's shape, the patch budget and the pixel normalization live in vision.json

, which the converter writes from the release's own nested vision_config

and from preprocessor_config.json

. K3 normalizes to [-1, 1] with mean = std = 0.5.

That last sentence was wrong here for a day, and the way it was wrong is worth keeping. This section used to say K3 ships no preprocessor config, so the normalization was "the CLIP convention this lineage of towers uses rather than a value read out of the release" โ€” an assumption, labelled as one. The release does ship the file; the down fetched a hardcoded list of filenames and never asked the repo what it contained. The tower still matched its oracle at 2.3e-06 throughout, because the oracle is fed random pixels and never touches the normalization. An honest caveat is not a substitute for reading the file.

build model-free suite backend
macOS arm64 yes 23 pass / 0 fail / 11 skip NEON
Linux arm64 yes 23 pass / 0 fail / 11 skip NEON
Linux x86_64 yes 23 pass / 0 fail / 11 skip AVX2
Windows x86_64 yes container, CLI and forward pass โ€” see below AVX2

The first three run the same suite and now agree check for check: same 23 passes, same 11 skips, same list. CI has no container, so tests/run.sh

builds a synthetic one and the checks that need real weights say SKIP rather than passing quietly. All three also pass the sanitizer suite and 400 fuzz cases.

Windows is cross-compiled with MinGW-w64 on a Linux runner and then run on a Windows one: the binary reads a synthetic container, opens it from the CLI, and produces the same logits token-by-token as it does in chunks. It is not the same suite โ€” tests/run.sh

is a bash script that rebuilds first, and the Windows job runs binaries it did not build โ€” so what is claimed is what that job checks and no more. Nobody has run it on a real container there.

The platform is the variable, the suite is not, and that is the point: both Linux targets produce the same continuation as macOS and pass engine matches the PyTorch oracle when given a container, so the numerics carry across architectures and compilers.

SIMD is selected at run time from CPUID, so a single x86 binary uses AVX-512 where it exists and AVX2 where it does not. Accelerator backends are build-time options. A Metal backend exists and is off by default because it is correct and 22% slower: this engine issues several hundred small dependent matvecs per token, the worst possible shape for an accelerator, and the CPU path already runs at the machine's memory bandwidth.

src/        the engine โ€” 6,000 lines of C, no dependencies
  model.c     forward pass, MoE routing, KDA and MLA layers
  ecache.c    bounded LFRU expert cache over the banks
  vision.c    the 27-layer ViT and the projector into text space
  image.c     a file on disk to the patch tensor the tower wants
  waste.c     the public API
  simd_*.c    per-ISA kernels, selected at run time
cli/        the CLI, a client of the public API
serve/      the OpenAI-compatible server, the other client
  xtml.py     K3's prompt format, ported from the release's encoding_k3.py
  regions.py  its replies, back into reasoning / answer / tool calls
  engine.py   libwaste through ctypes, and the request queue
  server.py   /v1/chat/completions and friends
tools/      conversion and validation (Python, never at run time)
docs/       format, engine, backends, and what was learned
tests/      34 checks, and a diff against a PyTorch oracle given a model
  serve/      149 more for the server, incl. a differential vs upstream
examples/   chat.json for K3 and ChatML, the format a container carries
third_party/ stb_image.h, the single vendored header โ€” see its README

docs/LEARNED.md is the one to read before contributing. It records what was measured, including the optimizations that were refuted โ€” index-layout blocking, a 3-bit trunk, GPU offload, per-expert bit allocation โ€” with the numbers that killed them.

The API is not frozen, as above. The rest is stated plainly too, because finding these out for yourself is worse than reading them here:

  • a container carries its chat format in chat.json

, and the converter can only fill it in for a model whose format has been transcribed from its reference encoder โ€” K3 today. Neither Kimi release distributes a template, so for anything else the CLI says so and continues raw rather than guessing a format, which would produce plausible wrong answers instead of visibly wrong ones. Kimi-Linear is in that position now; - AVX-512 compiles and is dispatched from CPUID, and has still never executed an instruction. This laptop is ARM and its x86 emulation is Rosetta, which reports AVX2 and leaves the ZMM state disabled in XCR0; the hosted x86 runner is an AMD EPYC 7763, which answers avx512f/bw/dq/vl: no

, so CI saysAVX2 as well โ€” on Linux and on Windows both. The workflow prints the runner's flags before every build, so the day a runner has them theSIMD backend matches the CPU baselinecheck becomes the confirmation without anyone arranging it; Windows builds and runs, on one toolchain and one CPU. MinGW-w64 x86_64, cross-compiled, withsrc/platform.h

holding the six calls that are not POSIX: the positional read, the aligned allocation, the CPU count, the file size andFILE_FLAG_NO_BUFFERING

for the cache-bypass open. MSVC is a different port and has not been attempted โ€” the sources use GNU C. ARM64 Windows is not built. Neither is the page-cache bypass proven under load there: CI confirms Windows grants it on the runner's filesystem, which is not the same as measuring a hit rate against a container that does not fit in RAM;the expert checksum is off unless you ask for it(--verify

), and the trunk has no checksum at all. The first is a decision โ€” 5% of throughput on every token, against a container that is usually fine โ€” and it means the default build of a rotted container still answers with whatever the damaged bytes decode to. Run--verify

once after copying a container, to establish that it arrived intact.tools/verify_container.py

does not stand in for that: it re-derives records against thesource weights, so it wants torch and the original checkpoint on disk, and it answers whether the conversion was right rather than whether the copy still is. The second is not a decision: the trunk and the codebooks have nothing to check against in the format, and nothing has been built in its place;- every expert in a container is at the same bit width. The non-uniform per-expert allocation the format was designed around is not coming: it was measured on both models rather than built, and the importance it would allocate against does not vary โ€” the value of the third bit spreads at most 1.15x between experts in a layer and 1.01x between layers, so the optimal allocator and a coin flip write the same container. The one signal that is not flat, routing frequency, buys disk footprint and almost no I/O, which is the resource that is actually scarce.docs/LEARNED.mdยง20 has the table and the one measurement that would revive it.

Apache 2.0 โ€” see LICENSE. Copyright 2026 SQLite Cloud, Inc.

โ”€โ”€ more in #large-language-models 4 stories ยท sorted by recency
โ”€โ”€ more on @kimi k3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain โ€” perfect for shipping the agent you just read about.

$git push zahid main
โ†’ Live at https://your-agent.zahid.host โœ“
Get free account โ†’ Pricing
from โ‚ฌ0/mo ยท no card required
LIVE [news/show-hn-run-full-kimโ€ฆ] indexed:0 read:26min 2026-07-30 ยท โ€”