4-bit GGUF Quality for MoE Models: Why Only 3B of 180B Params Fire, and How to Prove Parity A developer demonstrated that 4-bit GGUF quantization of Mixture-of-Experts models can match full-precision quality when sensitive tensors are kept at higher precision, reporting MMLU-Pro 87.65 for a mixed-precision UD-Q4_K_XL graft that equals the full-precision baseline within noise. The writeup explains that a 180B-parameter MoE activates only about 3B parameters per token, so the router, shared experts, attention projections, and embedding/output layers must stay precise while the rarely fired expert feed-forward weights can be pushed to 4-bit. It includes the llama.cpp commands to reproduce the parity check against a held-out benchmark. Mixture-of-Experts MoE models look enormous on disk, but only a small slice of the weights does work on any single token. A 180B-parameter MoE can activate roughly 3B parameters per forward pass. That sparsity is exactly why 4-bit GGUF quantization behaves so differently on MoE than on a dense model, and why a mixed-precision scheme such as the UD-Q4 K XL style graft can hold accuracy that a flat 4-bit cast would lose. The honest part of the story is not the compression ratio. It is verification. A quantized file that loads and produces fluent text is not proof of anything. You prove parity by scoring the quantized model on a held-out benchmark and comparing it to the full-precision baseline, number against number. In the example below, a 4-bit mixed graft reaches MMLU-Pro 87.65, equal to the full-precision baseline inside noise. This post shows the reasoning and the exact llama.cpp commands to reproduce that check yourself. A dense transformer runs every weight for every token. A Mixture-of-Experts transformer replaces the big feed-forward block in each layer with many smaller expert blocks plus a router. For each token the router picks a few experts often two and ignores the rest. The attention layers and the router still run every time, but the overwhelming majority of the feed-forward parameters sit idle on any given token. The practical consequence: the "180B" headline counts total parameters, while the compute and the active memory traffic track the much smaller active set, on the order of 3B. Two numbers describe an MoE and they answer different questions: Quantization touches storage, so it operates on the total. Quality, however, is decided token by token on the active path. This gap is the whole reason MoE quantization is its own topic. Two forces pull in opposite directions. First, MoE is more forgiving in aggregate. Averaged over a corpus, errors introduced into rarely used experts barely move the output, because those experts rarely fire. A dense model has no such luxury: every weight is on the hot path for every token, so quantization error accumulates everywhere at once. Second, MoE is more fragile in specific places. The router is tiny but decisive. If quantization noise flips a routing decision, the token is suddenly served by a different expert, and the error is not a small perturbation but a discrete change of computation. Shared or "always-on" experts, the attention projections, and the embedding and output layers are also on the hot path for every token, so error there behaves like dense-model error. So the right mental model is not "MoE tolerates 4-bit." It is "MoE has a small hot path that must stay precise and a large cold path that can be squeezed hard." A flat 4-bit cast ignores that structure and spends the same few bits on the router as on a seldom-touched expert. That is where quality quietly leaks. Unsloth's "UD" Unsloth Dynamic GGUF variants, and similar mixed-precision schemes, treat the model as a graft rather than a uniform block. The idea is simple to state: keep the sensitive, always-active tensors at higher precision and push the rarely active expert tensors down to 4-bit. The K XL suffix signals a K-quant layout with an extra-large share of high-precision tensors retained. In practice a graft like this tends to protect: Everything else, principally the large pool of expert feed-forward weights, goes to 4-bit. Because that pool is both the biggest contributor to file size and the least active per token, you capture most of the compression while leaving the hot path near full precision. The result is a file a fraction of the original size whose per-token behavior barely changed, which is the only kind of compression worth shipping to an edge device. The important caveat: which tensors you keep high is a model-specific choice, and vendors publish ready-made UD builds precisely so you do not have to guess. The reasoning above is the general principle, not a recipe to hand-tune. Here is the rule that matters: a quant is guilty until it passes a benchmark. Perplexity on a tiny sample and a few good-looking chat replies are not evidence. You need a discriminating, held-out benchmark scored identically on both the full-precision model and the quantized model, then you compare the scores. MMLU-Pro is a reasonable choice because it is broad, hard enough to separate models, and standard enough that the harness is well tested. The workflow is: build the quant, score the baseline, score the quant, and accept only if the gap is within benchmark noise. If you are starting from a full-precision GGUF and want to roll your own 4-bit file rather than download a vendor UD build , llama.cpp does the cast: Convert a Hugging Face checkpoint to a full-precision GGUF first if needed python llama.cpp/convert hf to gguf.py ./model-full \ --outfile model-f16.gguf --outtype f16 Quantize to a 4-bit K-quant. Q4 K M is the common balanced target. ./llama-quantize model-f16.gguf model-Q4 K M.gguf Q4 K M Inspect which tensors landed at which precision ./llama-gguf model-Q4 K M.gguf | grep -iE "type|ffn|attn|token embd" For MoE specifically, prefer a published mixed-precision build UD-Q4 K XL or equivalent when one exists, because the per-tensor precision map is the hard part and it has already been tuned for that architecture. Run the identical harness against the baseline and the quant. Keep every knob fixed: same prompt template, same number of shots, same decoding settings, same subset of questions. Baseline: full-precision model llama-perplexity --multiple-choice \ -m model-f16.gguf \ -bf mmlu-pro.bin \ --ctx-size 4096 2 &1 | tee baseline.log Candidate: 4-bit mixed graft llama-perplexity --multiple-choice \ -m model-UD-Q4 K XL.gguf \ -bf mmlu-pro.bin \ --ctx-size 4096 2 &1 | tee quant.log Compare the final accuracy lines grep -i "final" baseline.log quant.log baseline f16 MMLU-Pro: 87.6x UD-Q4 K XL 4-bit MMLU-Pro: 87.65 Equal inside noise is a pass. If instead the quant drops several points, do not ship it: either the graft protected the wrong tensors, the router drifted, or the target precision was too aggressive for this model. The benchmark told you something the demo never would. A few guardrails so the comparison stays honest: It shines when the total parameter count is what blocks you you cannot fit the file in RAM while the active path is small enough that the device can actually run it. That is the classic edge case: a big MoE that no consumer machine can hold at full precision becomes loadable at 4-bit because the cold experts shrink dramatically. It is the wrong call when the model is dense, when you need the last fraction of a point on a sensitive task, or when you have not budgeted time to verify. The compression is free; the confidence is not. Does a smaller GGUF always mean lower quality? No. File size tracks total parameters, but MoE quality is decided on the small active path. A mixed-precision graft can shrink the file a lot while keeping the active path near full precision, so a much smaller file can score the same on a held-out benchmark. What is the difference between Q4 K M and a UD-Q4 K XL style build? Q4 K M applies a mostly uniform 4-bit K-quant. A UD-Q4 K XL style build keeps more of the sensitive, always-active tensors embeddings, attention, router at higher precision and pushes only the rarely active experts to 4-bit, which protects quality on MoE at a similar overall size. Why not just read the perplexity number? Perplexity on a small sample is a weak proxy and can look fine while task accuracy slips. A discriminating multiple-choice benchmark like MMLU-Pro, scored identically on both models, is far harder to fool and tells you whether real task behavior survived. How close is close enough to call it parity? Within the benchmark's own run-to-run noise. Score each model at least twice. If the baseline-to-quant gap is smaller than the spread between two baseline runs, treat it as equal. Report the actual margin rather than a pass or fail label. Can I just download a vendor quant instead of making my own? Usually yes, and for MoE it is often the better choice, because the per-tensor precision map is model specific and tuning it is the hard part. Still run the verification step yourself: trust the build, but confirm the number on your own harness. Does this apply to dense models too? The verification workflow applies to everything. The sparsity argument does not: dense models run every weight on every token, so 4-bit error accumulates across the whole network and there is no large cold path to squeeze cheaply.