{"slug": "does-anyone-know-how-the-vericoding-benchmark-ran-glm-4-5-and-deepseek-v3-1-one", "title": "Does anyone know how the vericoding benchmark ran GLM-4.5 and DeepSeek V3.1? Asking because one of them looks like the non-thinking variant", "summary": "A developer working on the Beneficial AI Foundation's vericoding benchmark questions whether the open-weight models GLM-4.5 and DeepSeek V3.1 were run with their thinking modes enabled, noting that the DeepSeek entry (deepseek-chat-v3.1) is the non-thinking variant and that the repo does not record run configurations. The developer also reports that the benchmark's headline scores for Verus (44%) and Lean (27%) do not match recomputations from the shipped CSV (31.1% and 18.0%, respectively), while Dafny's 82% matches (83.1%).", "body_md": "I have been working in the Beneficial AI Foundation’s vericoding benchmark — the one where a theorem prover, not a test suite, decides whether the model’s code is right. I put the SPARK/Ada track into it, and I posted my own numbers here earlier.\n\nWhile pulling their shipped results apart for a comparison, something stopped me, and I cannot answer it from the repo.\n\nRestricted to the 26 `numpy_simple`\n\ntasks I was working with, in Dafny:\n\n| model | passed | weights |\n|---|---|---|\n| gpt-5-mini | 23/26 — 88.5% | closed |\n| gemini-2.5-pro | 23/26 — 88.5% | closed |\n| claude-sonnet-4 | 23/26 — 88.5% | closed |\n| claude-opus-4.1 | 23/26 — 88.5% | closed |\n| gpt-5 | 22/26 — 84.6% | closed |\n| glm-4.5 | 20/26 — 76.9% | open |\n| grok-code | 19/26 — 73.1% | closed |\n| gemini-2.5-flash | 19/26 — 73.1% | closed |\n| deepseek-chat-v3.1 | 15/26 — 57.7% | open |\n| all nine | 187/234 — 79.9% |\n\nThe two open-weight entries are bottom and sixth. Those are the two anyone here could actually self-host, so they are the two I care about most.\n\nHere is what I cannot resolve.\n\n**The DeepSeek entry is deepseek-chat-v3.1.** V3.1 is a hybrid — it has a thinking mode and a non-thinking mode, and\n\n`chat`\n\nis the one with reasoning off. It is being scored against gpt-5, gemini-2.5-pro and claude-opus-4.1, which are reasoning models running as reasoning models.**GLM-4.5 also has a thinking mode**, and I cannot find anything in the repo that says whether it was enabled.\n\nIf both were run with the thing that makes them competitive switched off, then two open-weight models have been published at scores that are not what they can do, on a table people have been quoting since as a fair fight.\n\nI want to be careful here: **I do not know that this is what happened.** I am asking because I cannot find the run configuration recorded anywhere, and the model names are the only evidence I have.\n\nWhat makes me want an answer rather than assume one is that the repo’s own headline numbers do not all reproduce from the results file it ships. Computing from their CSV:\n\n| language | their headline | computes to |\n|---|---|---|\n| Dafny | 82% | 83.1% ✓ |\n| Verus | 44% | 31.1% |\n| Lean | 27% | 18.0% |\n\nDafny lands. The other two do not. That is not an accusation of anything, but it does mean I would rather ask than infer.\n\nSo, two questions for anyone who knows more than me.\n\nIf someone has hardware to serve either of them at a sensible quantisation, the test is cheap and entirely public: the same 26 Dafny tasks, their verifier, thinking on. I have neither the hardware nor a reason to be trusted as the only person running it, so I would rather it were somebody else, or several of us.\n\nWorth saying that this cuts against my own post. I have been describing my 30B local result as landing above both open-weight entrants. If those two come out in the eighties when run properly, that line of mine is wrong and I would want to know now rather than after I have said it a few more times.", "url": "https://wpnews.pro/news/does-anyone-know-how-the-vericoding-benchmark-ran-glm-4-5-and-deepseek-v3-1-one", "canonical_source": "https://forum.level1techs.com/t/does-anyone-know-how-the-vericoding-benchmark-ran-glm-4-5-and-deepseek-v3-1-asking-because-one-of-them-looks-like-the-non-thinking-variant/254745#post_1", "published_at": "2026-09-01 05:13:18+00:00", "updated_at": "2026-09-01 05:22:53.444592+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-tools"], "entities": ["Beneficial AI Foundation", "GLM-4.5", "DeepSeek V3.1", "deepseek-chat-v3.1", "Dafny", "Verus", "Lean"], "alternates": {"html": "https://wpnews.pro/news/does-anyone-know-how-the-vericoding-benchmark-ran-glm-4-5-and-deepseek-v3-1-one", "markdown": "https://wpnews.pro/news/does-anyone-know-how-the-vericoding-benchmark-ran-glm-4-5-and-deepseek-v3-1-one.md", "text": "https://wpnews.pro/news/does-anyone-know-how-the-vericoding-benchmark-ran-glm-4-5-and-deepseek-v3-1-one.txt", "jsonld": "https://wpnews.pro/news/does-anyone-know-how-the-vericoding-benchmark-ran-glm-4-5-and-deepseek-v3-1-one.jsonld"}}