# Benchmark for LLM Generated UI

> Source: <https://www.openui.com/benchmarks>
> Published: 2026-08-27 19:50:09+00:00

30 models on OpenUI, plus a 6-model comparison across 3 generative UI formats

OpenUIBenchmarks

# Generative UI Benchmark

Model comparison

Provider-coloured dots with one Pareto frontier. Key references are labeled; hover any point for its name and values.

## Models**20 / 30**

**xAI** 1

**OpenAI** 3

**Anthropic** 4

**Moonshot** 1

**Meta** 1

**Alibaba** 4

**Zhipu** 1

**Thinking Machines** 2

**DeepSeek** 2

**Microsoft** 1

**Mistral** 1

**IBM** 1

**Liquid** 1

**InclusionAI** 1

## View chart data

### Model comparison data

| Model | Provider | Family | Valid | Cost per task | Pricing | Frontier (all 30) | Shown in chart |
|---|---|---|---|---|---|---|---|
| Grok 4.6 | xAI | — | 99.5% | $0.0185 | List price | Yes | Yes |
| GPT-5.6 Sol | OpenAI | GPT-5.6 | 99.5% | $0.0476 | List price | No | Yes |
| Claude Opus 4.8 | Anthropic | Claude | 98.9% | $0.0493 | List price | No | Yes |
| Gemini 3.7 Flash | Gemini Flash | 98.9% | $0.0100 | List price | Yes | Yes | |
| Claude Sonnet 5 | Anthropic | Claude | 98.4% | $0.0202 | List price | No | Yes |
| GPT-5.6 Terra | OpenAI | GPT-5.6 | 98.4% | $0.0241 | List price | No | Yes |
| Kimi K3 | Moonshot | — | 96.2% | $0.0333 | List price | No | Yes |
| Claude Opus 5 | Anthropic | Claude | 96.2% | $0.0774 | List price | No | Yes |
| Muse Spark 1.2 | Meta | — | 96.2% | $0.0157 | List price | No | Yes |
| Claude Sonnet 4.6 | Anthropic | Claude | 92.9% | $0.0450 | List price | No | Yes |
| Qwen3.8 2.4T | Alibaba | Qwen3.8 | 91.8% | $0.0170 | List price | No | Yes |
| GLM-5.3 | Zhipu | — | 90.8% | $0.0120 | List price | No | Yes |
| Inkling Small | Thinking Machines | Inkling | 88.6% | $0.0033 | List price | Yes | Yes |
| DeepSeek V4 Flash | DeepSeek | DeepSeek V4 | 85.8% | $0.0004 | List price | Yes | Yes |
| DeepSeek V4 Pro | DeepSeek | DeepSeek V4 | 84.2% | $0.0089 | List price | No | Yes |
| GPT-5.6 Luna | OpenAI | GPT-5.6 | 83.7% | $0.0022 | List price | No | Yes |
| Qwen3.8 27B | Alibaba | Qwen3.8 | 78.8% | $0.0054 | List price | No | Yes |
| Gemini 3.6 Flash | Gemini Flash | 78.8% | $0.0091 | List price | No | Yes | |
| Gemini 3.5 Flash-Lite | Gemini Flash | 78.3% | $0.0041 | List price | No | Yes | |
| Inkling | Thinking Machines | Inkling | 73.4% | $0.0080 | List price | No | Yes |
| Qwen3.6 27B | Alibaba | Qwen3.6 | 68.5% | $0.0063 | List price | No | No |
| Qwen3.6 35B-A3B | Alibaba | Qwen3.6 | 61.4% | $0.0017 | List price | No | No |
| Gemma 4 31B | Gemma 4 | 46.7% | $0.0009 | List price | No | No | |
| Phi-4 | Microsoft | — | 44.0% | $0.0004 | List price | No | No |
| Gemma 4 26B-A4B | Gemma 4 | 29.9% | $0.0007 | List price | No | No | |
| Ministral 8B | Mistral | — | 27.2% | $0.0009 | List price | No | No |
| Granite 4.1 8B | IBM | — | 14.7% | $0.0002 | List price | Yes | No |
| LFM 2.5 2.6B | Liquid | — | 3.3% | $0 / free | Free | Yes | No |
| DiffusionGemma 26B-A4B | — | 13.0% | Not comparable | Self-hosted | No | No | |
| Ling 3.0 Tiny | InclusionAI | — | 9.8% | Not comparable | Self-hosted | No | No |

All 30models remain in this table and the downloads, including models hidden by the chart’s default filter. Frontier membership here is computed over the full board, so it does not shift with the chart’s selection; the drawn line is the frontier of the models currently shown. Self-hosted cost is unknown, not zero. Filter state is preserved in this page’s URL. Focused benchmark: [language and model results](/benchmarks/language). Full dataset: [JSON](/benchmarks/data.json), [CSV](/benchmarks/data.csv), or [agent Markdown](/benchmarks/agent.md).

### Format comparison data

| Model | Format | Valid | Cost per pass |
|---|---|---|---|
| GPT-5.6 Sol | OpenUI | 99.5% | $3.08 |
| Claude Opus 4.8 | OpenUI | 98.9% | $2.27 |
| Kimi K3 | OpenUI | 96.2% | $1.53 |
| Qwen3.8 2.4T | OpenUI | 91.8% | $0.81 |
| Muse Spark 1.2 | OpenUI | 96.2% | $0.71 |
| Gemini 3.7 Flash | OpenUI | 98.9% | $0.42 |
| GPT-5.6 Sol | A2UI | 96.2% | $6.84 |
| Claude Opus 4.8 | A2UI | 99.5% | $5.85 |
| Kimi K3 | A2UI | 95.7% | $3.16 |
| Qwen3.8 2.4T | A2UI | 91.3% | $1.90 |
| Muse Spark 1.2 | A2UI | 96.7% | $1.38 |
| Gemini 3.7 Flash | A2UI | 94.0% | $0.95 |
| GPT-5.6 Sol | json-render | 82.6% | $6.30 |
| Claude Opus 4.8 | json-render | 87.5% | $4.88 |
| Kimi K3 | json-render | 71.2% | $2.91 |
| Qwen3.8 2.4T | json-render | 81.0% | $1.46 |
| Muse Spark 1.2 | json-render | 81.5% | $1.34 |
| Gemini 3.7 Flash | json-render | 92.9% | $0.85 |

Focused benchmark: [framework comparison](/benchmarks/framework). Download this comparison as [JSON](/benchmarks/framework/data.json), [CSV](/benchmarks/framework/data.csv), or [agent Markdown](/benchmarks/framework/agent.md).

## Headline results

### 99.9% render rate

Only 1 of 1,104 screens rendered blank.

### 96.9% valid

Parsed, rooted, and passed every structural check.

### 1.8 to 2.6× cheaper

Half the tokens: lower cost, faster streaming.

### 88% repaired in flight

A sanitizer model patches only the broken lines, via incremental editing.

## Blank screens vs. renders

Did the generation actually render?

1,104 runs per format across 6 models.

Counted as each SDK’s own renderer produced them.

OpenUI had 1 blank screen in 1,104 runs, compared with 37 for A2UI and 4 for json-render.

Render rate

Runs that rendered against runs that came back blank, out of 184 per model.

## View render data

| Model | Format | Rendered | Blank | Render rate |
|---|---|---|---|---|
| GPT-5.6 Sol | OpenUI | 184 | 0 | 100.0% |
| Claude Opus 4.8 | OpenUI | 184 | 0 | 100.0% |
| Kimi K3 | OpenUI | 184 | 0 | 100.0% |
| Gemini 3.7 Flash | OpenUI | 184 | 0 | 100.0% |
| Qwen3.8 2.4T | OpenUI | 183 | 1 | 99.5% |
| Muse Spark 1.2 | OpenUI | 184 | 0 | 100.0% |
| GPT-5.6 Sol | A2UI | 179 | 5 | 97.3% |
| Claude Opus 4.8 | A2UI | 184 | 0 | 100.0% |
| Kimi K3 | A2UI | 179 | 5 | 97.3% |
| Gemini 3.7 Flash | A2UI | 178 | 6 | 96.7% |
| Qwen3.8 2.4T | A2UI | 168 | 16 | 91.3% |
| Muse Spark 1.2 | A2UI | 179 | 5 | 97.3% |
| GPT-5.6 Sol | json-render | 184 | 0 | 100.0% |
| Claude Opus 4.8 | json-render | 184 | 0 | 100.0% |
| Kimi K3 | json-render | 184 | 0 | 100.0% |
| Gemini 3.7 Flash | json-render | 184 | 0 | 100.0% |
| Qwen3.8 2.4T | json-render | 180 | 4 | 97.8% |
| Muse Spark 1.2 | json-render | 184 | 0 | 100.0% |

## Structural validity vs. render success

Was the output valid, and did anything render?

Valid: every part it refers to exists, and every setting is a real one. Render success: something appeared at all. 1,104 runs per format.

OpenUI leads on both; A2UI and json-render each trade one off against the other.

Structural validity and render success

Share of 1,104 runs per format; longer is better. Scales start at 75% and 95%, not zero. Arrows compare with OpenUI.

## View validity and render data

| Format | Valid runs | Partial runs | Blank runs | Validity | Render success |
|---|---|---|---|---|---|
| OpenUI | 1070 | 33 | 1 | 96.9% | 99.9% |
| A2UI | 1055 | 12 | 37 | 95.6% | 96.6% |
| json-render | 914 | 186 | 4 | 82.8% | 99.6% |

**What counts as valid.** Parses, has a root, every reference resolves, nothing orphaned or invented, no missing or out-of-range props, not truncated, and at least as many components as the brief has requirements. That last one is a count floor, not a check that each requirement was addressed: this measures structure, not coverage.

## Structural validity by model

How often the component graph holds together

Passes only if nothing is left dangling, nothing is invented, and every setting is valid. 184 runs per model, per format.

OpenUI leads overall: 96.9% vs 95.6% for A2UI. The lead changes by model, with two ties.

Structural validity by model

Runs whose component graph parses, resolves and validates, judged by each format's own SDK.

| Model | OpenUI | A2UI | json-render |
|---|---|---|---|
| SolOpenAI | 99.5 | 96.2 | 82.6 |
| Claude Opus 4.8Anthropic | 98.9 | 99.5 | 87.5 |
| Kimi K3Moonshot | 96.2 | 95.7 | 71.2 |
| Gemini 3.7 FlashGoogle | 98.9 | 94.0 | 92.9 |
| Qwen3.8 2.4TAlibaba | 91.8 | 91.3 | 81.0 |
| Muse Spark 1.2Meta | 96.2 | 96.7 | 81.5 |
| Average | 96.9 | 95.6 | 82.8 |

## Token consumption and cost

Tokens and dollars for the same screens

46 screens,

priced at provider list prices.

OpenUI uses fewer tokens and costs 1.8–2.6× less across every priced model.

Token consumption

System prompt + output for one screen

## View token data

| Format | System prompt | Mean output | Combined |
|---|---|---|---|
| OpenUI | 5,031 | 1,362 | 6,393 |
| A2UI | 12,610 | 2,823 | 15,433 |
| json-render | 7,651 | 3,258 | 10,909 |

Cost of one benchmark pass

The same 46 screens at list prices

| Model | OpenUI | A2UI | json-render |
|---|---|---|---|
| Gemini 3.7 FlashGoogle | $0.42 | $0.95↑2.3× | $0.85↑2.0× |
| Muse Spark 1.2Meta | $0.71 | $1.38↑1.9× | $1.34↑1.9× |
| Qwen3.8 2.4TAlibaba | $0.81 | $1.90↑2.3× | $1.46↑1.8× |
| Kimi K3Moonshot | $1.53 | $3.16↑2.1× | $2.91↑1.9× |
| Claude Opus 4.8Anthropic | $2.27 | $5.85↑2.6× | $4.88↑2.1× |
| SolOpenAI | $3.08 | $6.84↑2.2× | $6.30↑2.0× |

## Streaming speed

How long a screen takes to render

Mean output per screen,

decoded at 50 tokens per second.

A screen streams in about half the time, because there is about half as much to write.

Time to stream a screen

Mean output per screen, decoded at 50 tokens per second.

## View streaming data

| Format | Mean output tokens | Decode rate | Estimated seconds |
|---|---|---|---|
| OpenUI | 1,362 | 50 tok/s | 27.2s |
| A2UI | 2,823 | 50 tok/s | 56.5s |
| json-render | 3,258 | 50 tok/s | 65.2s |

## Structural validity by screen complexity

Validity as requirements increase

46 briefs across 5 complexity bands,

about 9 per band.

As screens get harder, OpenUI and A2UI stay above 90%; json-render falls to 72%.

Structural validity by screen complexity

The more a brief asks for, the less every format delivers. OpenUI and A2UI track each other closely at every size; json-render falls fastest and furthest.

## View complexity data

| Requirements per screen | Runs | OpenUI | A2UI | json-render |
|---|---|---|---|---|
| 2–3 | 240 | 100.0% | 99.6% | 89.6% |
| 4–6 | 240 | 99.6% | 98.8% | 81.7% |
| 7–9 | 240 | 98.8% | 94.2% | 85.4% |
| 11–13 | 192 | 91.1% | 91.7% | 68.2% |
| 16–18 | 192 | 90.6% | 93.2% | 71.9% |

## Production repair

What gets fixed before users see it

Based on real production data:

OpenUI Cloud traffic, not benchmark runs.

Only 0.9% of generations reach a user broken. Of the ones that fail validation, 88% are repaired via incremental editing.

Repair funnel

What happens to a broken generation before it can reach a user.

**100%** All generationsNo model callParser fixes syntax issues first

**7%** Fail validationNo model callStructural issues like dangling refs and bad enums

**0.9%** Reach a user brokenOne model call88% are repaired via incremental editing

## View repair data

| Stage | Share of generations | What happens | Model call |
|---|---|---|---|
| All generations | 100% | Parser fixes syntax issues first | No model call |
| Fail validation | 7% | Structural issues like dangling refs and bad enums | No model call |
| Reach a user broken | 0.9% | 88% are repaired via incremental editing | One model call |

Production traffic, not benchmark runs.
