{"slug": "show-hn-a-genai-image-benchmark-for-production-design-work", "title": "Show HN: A GenAI image benchmark for production design work", "summary": "A new benchmark for generative AI image models, built for production design work, scores 15 models across 37 canonical prompts and 1,662 generated images, finding that 13 of 15 models fail at transparency, no model hits exact brand colors (best score 3.67 of 5 on CIEDE2000), and rendered text is often not shippable. The benchmark, last run on Jul 6, 2026, measures typography, brand-color fidelity, transparency, composition, consistency, cost, and latency, and offers a 2026 Benchmark Report PDF.", "body_md": "# Benchmark GenAI image models\n\nThe same prompt, every model, side by side. We run identical prompts across the leading image generation models and score the results on what production design work actually needs: typography, brand-color fidelity, transparency, composition, consistency, cost and latency. When a new model drops, the whole suite re-runs.\n\n15\n\nModels benchmarked\n\n37\n\nCanonical prompts\n\n1662\n\nImages generated\n\nJul 6, 2026\n\nLast benchmark run\n\nLogo wordmark for a company called \"Aurelia\", geometric sans-serif, flat black on white, perfectly spelled, centered.\n\n## What are you building?\n\nStart from the job, not the model. Each path weights the same measurements for what actually decides success in that use case.\n\n## Same prompt. Every model. Side by side.\n\nIdentical text, default parameters, fixed seeds, no per-model tuning. Every result ships with its seed, latency, cost and native resolution, so you compare models, not prompt engineering.\n\n[See the full comparison and scores](ai-benchmarks/prompts/t04-wordmark/)\n\n### Which model should you build on?\n\nEvery model side by side: cost per image, latency, native resolution and per-criterion scores, with a full result gallery and per-category breakdown on every model profile.\n\n### The canonical prompt suite\n\nTypography, vector style, brand color, transparency, composition and spatial adherence: each prompt stresses a criterion that decides whether an asset ships, and each has a result grid across all models.\n\n## What the data says\n\nThe findings that matter if you are integrating AI imagery into a product: where models fall short of production requirements, measured.\n\n### 13 of 15 models fail at transparency\n\nWe asked every model for transparent PNGs and measured the alpha channel of what came back. Almost none of it survives contact with a real sticker, merch or cut-out pipeline.\n\n### No model hits your exact hex\n\nMeasured with CIEDE2000 against required brand colors, the best model scores 3.67 of 5 and most of the field lands below 3. Close-enough color is not the color in your brand book.\n\n### Rendered text: close is not shippable\n\nEven the best model occasionally breaks a headline, and the model famous for text lands mid-field. One wrong character means regenerating the whole image, unless the words are editable layers.\n\n### Get the 2026 Benchmark Report\n\n15 models, 37 prompts, 1,662 measured images. Every finding and ranking from this benchmark, with the methodology behind the numbers, as a PDF in your inbox.\n\nCheck your inbox.\n\nYour report is on the way.\n\n## How we score, and why you can trust it\n\nEvery criterion is scored by the cheapest tier that is reliable for it. Numbers a machine can measure are measured; judgments that need eyes get them. Methodology is published in full, models get zero special treatment, and scores are never silently restated.\n\n### Measured\n\nResolution, latency, cost, alpha-channel quality and brand-color drift (CIEDE2000 against the requested hex values): computed from every generation, reproducible from the recorded originals.\n\n### Judged\n\nPrompt adherence checklists and composition checks run through a vision-language judge, calibrated against the expert panel. Pending in the pilot dataset and always labeled as such.\n\n### Rated\n\nDesign-readiness and aesthetics come from a blind expert panel: model names hidden, three raters per cell, agreement reported. Arrives with the frozen suite.", "url": "https://wpnews.pro/news/show-hn-a-genai-image-benchmark-for-production-design-work", "canonical_source": "https://img.ly/ai-benchmarks/", "published_at": "2026-08-04 08:26:29+00:00", "updated_at": "2026-08-04 08:52:36.840508+00:00", "lang": "en", "topics": ["generative-ai", "ai-tools", "ai-research"], "entities": ["Aurelia", "CIEDE2000"], "alternates": {"html": "https://wpnews.pro/news/show-hn-a-genai-image-benchmark-for-production-design-work", "markdown": "https://wpnews.pro/news/show-hn-a-genai-image-benchmark-for-production-design-work.md", "text": "https://wpnews.pro/news/show-hn-a-genai-image-benchmark-for-production-design-work.txt", "jsonld": "https://wpnews.pro/news/show-hn-a-genai-image-benchmark-for-production-design-work.jsonld"}}