cd /news/artificial-intelligence/benchmark-for-llm-generated-ui · home topics artificial-intelligence article
[ARTICLE · art-113490] src=openui.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Benchmark for LLM Generated UI

OpenUI's generative UI benchmark of 30 models found xAI's Grok 4.6 and OpenAI's GPT-5.6 Sol tied for the highest validity at 99.5%, with Grok 4.6 the cheapest frontier model at $0.0185 per task. Across three generative UI formats, GPT-5.6 Sol led OpenUI at 99.5% validity, while Claude Opus 4.8 topped A2UI at 99.5%.

read11 min views1 publishedAug 27, 2026

30 models on OpenUI, plus a 6-model comparison across 3 generative UI formats

OpenUIBenchmarks

Model comparison

Provider-coloured dots with one Pareto frontier. Key references are labeled; hover any point for its name and values.

Models20 / 30 #

xAI 1

OpenAI 3

Anthropic 4

Moonshot 1

Meta 1

Alibaba 4

Zhipu 1

Thinking Machines 2

DeepSeek 2

Microsoft 1

Mistral 1

IBM 1

Liquid 1

InclusionAI 1

View chart data #

Model comparison data

Model Provider Family Valid Cost per task Pricing Frontier (all 30) Shown in chart
Grok 4.6 xAI 99.5% $0.0185 List price Yes Yes
GPT-5.6 Sol OpenAI GPT-5.6 99.5% $0.0476 List price No Yes
Claude Opus 4.8 Anthropic Claude 98.9% $0.0493 List price No Yes
Gemini 3.7 Flash Gemini Flash 98.9% $0.0100 List price Yes Yes
Claude Sonnet 5 Anthropic Claude 98.4% $0.0202 List price No Yes
GPT-5.6 Terra OpenAI GPT-5.6 98.4% $0.0241 List price No Yes
Kimi K3 Moonshot 96.2% $0.0333 List price No Yes
Claude Opus 5 Anthropic Claude 96.2% $0.0774 List price No Yes
Muse Spark 1.2 Meta 96.2% $0.0157 List price No Yes
Claude Sonnet 4.6 Anthropic Claude 92.9% $0.0450 List price No Yes
Qwen3.8 2.4T Alibaba Qwen3.8 91.8% $0.0170 List price No Yes
GLM-5.3 Zhipu 90.8% $0.0120 List price No Yes
Inkling Small Thinking Machines Inkling 88.6% $0.0033 List price Yes Yes
DeepSeek V4 Flash DeepSeek DeepSeek V4 85.8% $0.0004 List price Yes Yes
DeepSeek V4 Pro DeepSeek DeepSeek V4 84.2% $0.0089 List price No Yes
GPT-5.6 Luna OpenAI GPT-5.6 83.7% $0.0022 List price No Yes
Qwen3.8 27B Alibaba Qwen3.8 78.8% $0.0054 List price No Yes
Gemini 3.6 Flash Gemini Flash 78.8% $0.0091 List price No Yes
Gemini 3.5 Flash-Lite Gemini Flash 78.3% $0.0041 List price No Yes
Inkling Thinking Machines Inkling 73.4% $0.0080 List price No Yes
Qwen3.6 27B Alibaba Qwen3.6 68.5% $0.0063 List price No No
Qwen3.6 35B-A3B Alibaba Qwen3.6 61.4% $0.0017 List price No No
Gemma 4 31B Gemma 4 46.7% $0.0009 List price No No
Phi-4 Microsoft 44.0% $0.0004 List price No No
Gemma 4 26B-A4B Gemma 4 29.9% $0.0007 List price No No
Ministral 8B Mistral 27.2% $0.0009 List price No No
Granite 4.1 8B IBM 14.7% $0.0002 List price Yes No
LFM 2.5 2.6B Liquid 3.3% $0 / free Free Yes No
DiffusionGemma 26B-A4B 13.0% Not comparable Self-hosted No No
Ling 3.0 Tiny InclusionAI 9.8% Not comparable Self-hosted No No

All 30models remain in this table and the downloads, including models hidden by the chart’s default filter. Frontier membership here is computed over the full board, so it does not shift with the chart’s selection; the drawn line is the frontier of the models currently shown. Self-hosted cost is unknown, not zero. Filter state is preserved in this page’s URL. Focused benchmark: language and model results. Full dataset: JSON, CSV, or agent Markdown.

Format comparison data

Model Format Valid Cost per pass
GPT-5.6 Sol OpenUI 99.5% $3.08
Claude Opus 4.8 OpenUI 98.9% $2.27
Kimi K3 OpenUI 96.2% $1.53
Qwen3.8 2.4T OpenUI 91.8% $0.81
Muse Spark 1.2 OpenUI 96.2% $0.71
Gemini 3.7 Flash OpenUI 98.9% $0.42
GPT-5.6 Sol A2UI 96.2% $6.84
Claude Opus 4.8 A2UI 99.5% $5.85
Kimi K3 A2UI 95.7% $3.16
Qwen3.8 2.4T A2UI 91.3% $1.90
Muse Spark 1.2 A2UI 96.7% $1.38
Gemini 3.7 Flash A2UI 94.0% $0.95
GPT-5.6 Sol json-render 82.6% $6.30
Claude Opus 4.8 json-render 87.5% $4.88
Kimi K3 json-render 71.2% $2.91
Qwen3.8 2.4T json-render 81.0% $1.46
Muse Spark 1.2 json-render 81.5% $1.34
Gemini 3.7 Flash json-render 92.9% $0.85

Focused benchmark: framework comparison. Download this comparison as JSON, CSV, or agent Markdown.

Headline results #

99.9% render rate

Only 1 of 1,104 screens rendered blank.

96.9% valid

Parsed, rooted, and passed every structural check.

1.8 to 2.6× cheaper

Half the tokens: lower cost, faster streaming.

88% repaired in flight

A sanitizer model patches only the broken lines, via incremental editing.

Blank screens vs. renders #

Did the generation actually render?

1,104 runs per format across 6 models.

Counted as each SDK’s own renderer produced them.

OpenUI had 1 blank screen in 1,104 runs, compared with 37 for A2UI and 4 for json-render.

Render rate

Runs that rendered against runs that came back blank, out of 184 per model.

View render data #

Model Format Rendered Blank Render rate
GPT-5.6 Sol OpenUI 184 0 100.0%
Claude Opus 4.8 OpenUI 184 0 100.0%
Kimi K3 OpenUI 184 0 100.0%
Gemini 3.7 Flash OpenUI 184 0 100.0%
Qwen3.8 2.4T OpenUI 183 1 99.5%
Muse Spark 1.2 OpenUI 184 0 100.0%
GPT-5.6 Sol A2UI 179 5 97.3%
Claude Opus 4.8 A2UI 184 0 100.0%
Kimi K3 A2UI 179 5 97.3%
Gemini 3.7 Flash A2UI 178 6 96.7%
Qwen3.8 2.4T A2UI 168 16 91.3%
Muse Spark 1.2 A2UI 179 5 97.3%
GPT-5.6 Sol json-render 184 0 100.0%
Claude Opus 4.8 json-render 184 0 100.0%
Kimi K3 json-render 184 0 100.0%
Gemini 3.7 Flash json-render 184 0 100.0%
Qwen3.8 2.4T json-render 180 4 97.8%
Muse Spark 1.2 json-render 184 0 100.0%

Structural validity vs. render success #

Was the output valid, and did anything render?

Valid: every part it refers to exists, and every setting is a real one. Render success: something appeared at all. 1,104 runs per format.

OpenUI leads on both; A2UI and json-render each trade one off against the other.

Structural validity and render success

Share of 1,104 runs per format; longer is better. Scales start at 75% and 95%, not zero. Arrows compare with OpenUI.

View validity and render data #

Format Valid runs Partial runs Blank runs Validity Render success
OpenUI 1070 33 1 96.9% 99.9%
A2UI 1055 12 37 95.6% 96.6%
json-render 914 186 4 82.8% 99.6%

What counts as valid. Parses, has a root, every reference resolves, nothing orphaned or invented, no missing or out-of-range props, not truncated, and at least as many components as the brief has requirements. That last one is a count floor, not a check that each requirement was addressed: this measures structure, not coverage.

Structural validity by model #

How often the component graph holds together

Passes only if nothing is left dangling, nothing is invented, and every setting is valid. 184 runs per model, per format.

OpenUI leads overall: 96.9% vs 95.6% for A2UI. The lead changes by model, with two ties.

Structural validity by model

Runs whose component graph parses, resolves and validates, judged by each format's own SDK.

Model OpenUI A2UI json-render
SolOpenAI 99.5 96.2 82.6
Claude Opus 4.8Anthropic 98.9 99.5 87.5
Kimi K3Moonshot 96.2 95.7 71.2
Gemini 3.7 FlashGoogle 98.9 94.0 92.9
Qwen3.8 2.4TAlibaba 91.8 91.3 81.0
Muse Spark 1.2Meta 96.2 96.7 81.5
Average 96.9 95.6 82.8

Token consumption and cost #

Tokens and dollars for the same screens

46 screens,

priced at provider list prices.

OpenUI uses fewer tokens and costs 1.8–2.6× less across every priced model.

Token consumption

System prompt + output for one screen

View token data #

Format System prompt Mean output Combined
OpenUI 5,031 1,362 6,393
A2UI 12,610 2,823 15,433
json-render 7,651 3,258 10,909

Cost of one benchmark pass

The same 46 screens at list prices

Model OpenUI A2UI json-render
Gemini 3.7 FlashGoogle $0.42 $0.95↑2.3× $0.85↑2.0×
Muse Spark 1.2Meta $0.71 $1.38↑1.9× $1.34↑1.9×
Qwen3.8 2.4TAlibaba $0.81 $1.90↑2.3× $1.46↑1.8×
Kimi K3Moonshot $1.53 $3.16↑2.1× $2.91↑1.9×
Claude Opus 4.8Anthropic $2.27 $5.85↑2.6× $4.88↑2.1×
SolOpenAI $3.08 $6.84↑2.2× $6.30↑2.0×

Streaming speed #

How long a screen takes to render

Mean output per screen,

decoded at 50 tokens per second.

A screen streams in about half the time, because there is about half as much to write.

Time to stream a screen

Mean output per screen, decoded at 50 tokens per second.

View streaming data #

Format Mean output tokens Decode rate Estimated seconds
OpenUI 1,362 50 tok/s 27.2s
A2UI 2,823 50 tok/s 56.5s
json-render 3,258 50 tok/s 65.2s

Structural validity by screen complexity #

Validity as requirements increase

46 briefs across 5 complexity bands,

about 9 per band.

As screens get harder, OpenUI and A2UI stay above 90%; json-render falls to 72%.

Structural validity by screen complexity

The more a brief asks for, the less every format delivers. OpenUI and A2UI track each other closely at every size; json-render falls fastest and furthest.

View complexity data #

Requirements per screen Runs OpenUI A2UI json-render
2–3 240 100.0% 99.6% 89.6%
4–6 240 99.6% 98.8% 81.7%
7–9 240 98.8% 94.2% 85.4%
11–13 192 91.1% 91.7% 68.2%
16–18 192 90.6% 93.2% 71.9%

Production repair #

What gets fixed before users see it

Based on real production data:

OpenUI Cloud traffic, not benchmark runs.

Only 0.9% of generations reach a user broken. Of the ones that fail validation, 88% are repaired via incremental editing.

Repair funnel

What happens to a broken generation before it can reach a user.

100% All generationsNo model callParser fixes syntax issues first

7% Fail validationNo model callStructural issues like dangling refs and bad enums

0.9% Reach a user brokenOne model call88% are repaired via incremental editing

View repair data #

Stage Share of generations What happens Model call
All generations 100% Parser fixes syntax issues first No model call
Fail validation 7% Structural issues like dangling refs and bad enums No model call
Reach a user broken 0.9% 88% are repaired via incremental editing One model call

Production traffic, not benchmark runs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openui 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/benchmark-for-llm-ge…] indexed:0 read:11min 2026-08-27 ·