{"slug": "veriloop-e2-release-27b-post-trained-model-and-full-gguf-precision-ladder-from-m", "title": "VeriLoop E2 Release: 27B Post-Trained Model and Full GGUF Precision Ladder from BF16 to IQ1_M", "summary": "A small independent test of the VeriLoop E2 GGUF release found that the IQ1_M and Q4_K_M quantizations of the 27B post-trained model made identical decisions on all 8 forced-choice prompts, including the same 3 incorrect answers, when run on an Nvidia L4. The tester framed the 8/8 decision agreement as a narrow behavioral sanity check that complements the release's published PPL/KLD distribution-retention measurements, and explicitly not as evidence that IQ1_M preserves free-form reasoning, agentic behavior, or downstream benchmark performance. The VeriLoop E2 release, published by Tsinghua SIGS Robot Lab on Hugging Face, ships a full GGUF precision ladder from BF16 down to IQ1_M alongside separate quantization-retention and downstream benchmark retention evidence.", "body_md": "For now, I tried a quick test:\n\nThe GGUF side of this release looked especially interesting to me because you already separate **quantization-retention measurements** from **downstream benchmark retention** in the announcement.\n\nSo I tried a small independent sanity check on the official [VeriLoop E2 GGUF release](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF), comparing **IQ1_M vs Q4_K_M** on an L4.\n\nThe result was simple:\n\n|  | IQ1_M | Q4_K_M | \n|---|---|---|\n| Parsed decisions | 8/8 | 8/8 | \n| Correct on my tiny panel | 5/8 | 5/8 | \n| Decision agreement | - | **8/8** | \n\nSo, on this very small forced-choice panel, the two official quants made **exactly the same eight decisions**, including the same three wrong ones.\n\nI would interpret this narrowly: it is a small behavioral sanity check complementary to the PPL/KLD measurements in the release, **not** evidence that IQ1_M preserves free-form reasoning, agentic behavior, or downstream benchmark performance in general.\n\nStill, given how aggressive the footprint reduction is, I thought the 8/8 agreement was worth reporting.\n\nExact setup and what I think this does/does not show\nOverall, the part I found most encouraging is that the release already exposes several different evidence layers instead of collapsing everything into a single headline score:\n\n```\ncheckpoint\ntraining method\nHarness / verifier\nbenchmark evidence\nGGUF quantization\ndistribution-retention measurements\nruntime validation\n```\n\nMy small IQ1_M/Q4_K_M check adds only one tiny additional point to that map, but so far it points in the same direction as the published GGUF retention measurements.\n\nFor this kind of release, I think keeping those layers separate is more useful than trying to turn all of them into one “quality” number.\n\nThanks for publishing the GGUF ladder and the evidence alongside it — having enough public surface to independently check even a small slice is useful.", "url": "https://wpnews.pro/news/veriloop-e2-release-27b-post-trained-model-and-full-gguf-precision-ladder-from-m", "canonical_source": "https://discuss.huggingface.co/t/veriloop-e2-release-27b-post-trained-model-and-full-gguf-precision-ladder-from-bf16-to-iq1-m/180723#post_2", "published_at": "2026-09-27 01:03:13+00:00", "updated_at": "2026-09-27 01:31:27.590252+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-infrastructure", "mlops"], "entities": ["VeriLoop E2", "Tsinghua SIGS Robot Lab", "Hugging Face", "GGUF", "IQ1_M", "Q4_K_M", "BF16", "Nvidia L4"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/veriloop-e2-release-27b-post-trained-model-and-full-gguf-precision-ladder-from-m", "markdown": "https://wpnews.pro/news/veriloop-e2-release-27b-post-trained-model-and-full-gguf-precision-ladder-from-m.md", "text": "https://wpnews.pro/news/veriloop-e2-release-27b-post-trained-model-and-full-gguf-precision-ladder-from-m.txt", "jsonld": "https://wpnews.pro/news/veriloop-e2-release-27b-post-trained-model-and-full-gguf-precision-ladder-from-m.jsonld"}}