cd /news/artificial-intelligence/comparing-open-weight-models-vs-clos… · home topics artificial-intelligence article
[ARTICLE · art-121059] src=pub.towardsai.net ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Comparing Open-Weight Models vs Closed Frontier Models: Debunking Some of the Hype

As of 2026, open-weight AI models have nearly closed the intelligence gap with closed US frontier models, with China's Kimi K3 (Moonshot AI) and GLM-5.3 (Z.ai) scoring 60 on the Artificial Analysis Intelligence index, just 3 points behind the top closed model Claude Opus 5 (max) at 63. However, UK AISI and US CAISI joint evaluations show Kimi K3 lags significantly on cybersecurity, scoring 32% on ExploitBench versus 76.2% for leading US frontier models, and DeepSeek V4's capabilities trail the frontier by about 8 months. The article advises that model choice should depend on context fit, data protection, and total cost, not just hype.

read4 min views1 publishedSep 4, 2026

The two charts explain most of what is happening in 2026.

The first: On the Artificial Analysis* Intelligence index, the top open-weight model is now within a couple of points of the best closed US model. Here are some details:

Closed US frontier: Claude Opus 5 (max) at 63, Claude Fable 5 at 62, and GPT-5.6 Sol (max) at 61, the top three overall.

Open weights, top three: Kimi K3 (max by Moonshot AI) at 60, GLM-5.3 (max by Z.ai) at 60, Qwen3.8 2.4T (Alibaba) at 58. The second chart breaks open weights down by country: China AI labs hold the top of the open-weight board (with Kimi K3 at 60). The best US open-weight is NVIDIA’s Nemotron 3 Ultra at 48, the highest ever for a US lab.

The “catching up” debate is almost over now. What matters more is a different decision: which model correctly understands your context and fits perfectly, or near perfectly, inside your AI harness. Of course, if data protection is your top priority, that will steer you toward the model you ultimately choose.

Now let’s closely compare a few things that will help you decide, and maybe not just follow the hype.

  1. The intelligence gap: widest on cybersecurity

Kimi K3 performs significantly below the most recent frontier cyber-capable models on preliminary cyber evaluations run by UK AISI* and CAISI*.

UK AISI and US CAISI ran a joint preliminary evaluation of Kimi K3’s cyber capabilities in July 2026. On ExploitBench (developing working exploits), K3 scored 32% against 76.2% for the leading US frontier models.

DeepSeek V4 is more cost efficient than other models of similar capability. However, according to CAISI’s evaluations, DeepSeek V4’s capabilities lag behind the frontier by about 8 months.

Kimi K3 holds #1 on LMArena’s Frontend Code Arena and #1 on AA’s AutomationBench-AA* (i.e. handling automated tasks) with 53%, beating the US frontier outright in those arenas. On the Vals coding index* it effectively ties with Opus 5 (74.70% vs 74.82%). On AA-Omniscience*, Kimi K3 answers correctly 46% of the time, but when it is wrong, it hallucinates (guesses confidently) rather than admits it does not know 51% of the time. Its output speed (39 tokens/sec) is also about two-thirds of the Opus class.

  1. Price: an order of magnitude, per token

Official list prices, per million tokens (as per August, 2026):

At the top of the open-weight market, the price gap has almost closed.

Kimi K3’s $3/$15 sits within 25 to 33 % of GPT-5.6 Sol’s new rate, and is now identical to Claude Sonnet 5’s standard pricing.

However, note that cheaper models sometimes need longer reasoning, more retries, or always a human in the loop. A premium model that succeeds in one pass can be the cheaper alternative sometimes.

  1. “Open-weight” means more than just one thing

Size: Kimi K3 totals 2.8T parameters, with an MXFP4 checkpoint (used for reducing memory requirements) around 1.5 TB to host. DeepSeek V4 Pro totals 1.6T. GLM-5.3 comes in near 750B (about 740 GB at FP8). Mixture-of-experts architecture means only a fraction activates per token: Kimi activates 16 of 896 experts, DeepSeek V4 Pro about 49B of its 1.6T. This is why these models are cheap to call and unrealistic to self-host casually. “Open” does not mean “runs on your laptop.”

License: GLM-5.2 was shipped under a genuine MIT license. Its successor did not: GLM-5.3’s weights landed on Hugging Face on August 28 under a new custom license, reported to carry restrictions for the largest companies (those with $10B+ in revenue). Kimi K3 ships a modified MIT “Kimi K3 License,” free to self-host, but model-as-a-service providers above $20M in annual revenue need a separate agreement with Moonshot.

Control: Self-hosting is the real deal no doubt. Your hardware, your data, no per-token fees, no vendor lock-in. However, self hosting actually becomes an infrastructure project.

  1. Compliance

SOC 2 reports, HIPAA BAAs, and GDPR data agreements are standard on US enterprise plans and rare on hosted Chinese open-weight APIs.

US frontier labs still lead on enterprise tooling: SLAs, admin controls, safety settings, vendor accountability.

So, how do you decide?

This study aggregates published evaluations from Artificial Analysis, CAISI, UK AISI, and vendor system cards, so you can check every source and apply your own judgment.

Metrics differ across evaluators, and there is no enterprise standardized way to compare models today. Treat every number here as one data point. This is not a verdict.

I write about AI architecture, security, and open-weight models and harnesses. If this is usefully for you, share it with others.

* Artificial Analysis: an independent research group that benchmarks and ranks AI models across many tasks.

* AA’s AutomationBench-AA: an Artificial Analysis benchmark testing how well a model handles real, automated work tasks.

* Vals coding index: a benchmark that scores how well AI models write and fix code.

* AA-Omniscience: an Artificial Analysis benchmark measuring how often a model confidently states false information (its hallucination rate).

* UK AISI: the UK government’s AI Security Institute, which tests AI models for safety and security risks.

* CAISI: the US Center for AI Standards and Innovation (under NIST), which evaluates AI model capabilities and risks.
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @moonshot ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/comparing-open-weigh…] indexed:0 read:4min 2026-09-04 ·