# Show HN: A GenAI image benchmark for production design work

> Source: <https://img.ly/ai-benchmarks/>
> Published: 2026-08-04 08:26:29+00:00

# Benchmark GenAI image models

The same prompt, every model, side by side. We run identical prompts across the leading image generation models and score the results on what production design work actually needs: typography, brand-color fidelity, transparency, composition, consistency, cost and latency. When a new model drops, the whole suite re-runs.

15

Models benchmarked

37

Canonical prompts

1662

Images generated

Jul 6, 2026

Last benchmark run

Logo wordmark for a company called "Aurelia", geometric sans-serif, flat black on white, perfectly spelled, centered.

## What are you building?

Start from the job, not the model. Each path weights the same measurements for what actually decides success in that use case.

## Same prompt. Every model. Side by side.

Identical text, default parameters, fixed seeds, no per-model tuning. Every result ships with its seed, latency, cost and native resolution, so you compare models, not prompt engineering.

[See the full comparison and scores](ai-benchmarks/prompts/t04-wordmark/)

### Which model should you build on?

Every model side by side: cost per image, latency, native resolution and per-criterion scores, with a full result gallery and per-category breakdown on every model profile.

### The canonical prompt suite

Typography, vector style, brand color, transparency, composition and spatial adherence: each prompt stresses a criterion that decides whether an asset ships, and each has a result grid across all models.

## What the data says

The findings that matter if you are integrating AI imagery into a product: where models fall short of production requirements, measured.

### 13 of 15 models fail at transparency

We asked every model for transparent PNGs and measured the alpha channel of what came back. Almost none of it survives contact with a real sticker, merch or cut-out pipeline.

### No model hits your exact hex

Measured with CIEDE2000 against required brand colors, the best model scores 3.67 of 5 and most of the field lands below 3. Close-enough color is not the color in your brand book.

### Rendered text: close is not shippable

Even the best model occasionally breaks a headline, and the model famous for text lands mid-field. One wrong character means regenerating the whole image, unless the words are editable layers.

### Get the 2026 Benchmark Report

15 models, 37 prompts, 1,662 measured images. Every finding and ranking from this benchmark, with the methodology behind the numbers, as a PDF in your inbox.

Check your inbox.

Your report is on the way.

## How we score, and why you can trust it

Every criterion is scored by the cheapest tier that is reliable for it. Numbers a machine can measure are measured; judgments that need eyes get them. Methodology is published in full, models get zero special treatment, and scores are never silently restated.

### Measured

Resolution, latency, cost, alpha-channel quality and brand-color drift (CIEDE2000 against the requested hex values): computed from every generation, reproducible from the recorded originals.

### Judged

Prompt adherence checklists and composition checks run through a vision-language judge, calibrated against the expert panel. Pending in the pilot dataset and always labeled as such.

### Rated

Design-readiness and aesthetics come from a blind expert panel: model names hidden, three raters per cell, agreement reported. Arrives with the frozen suite.
