{"slug": "the-industry-trained-an-ai-design-monoculture-people-preferred-the-models-that", "title": "The industry trained an AI design monoculture. People preferred the models that broke from it.", "summary": "HumanSignal's blind round-robin design study, run in Label Studio with 4,500 votes from 29 annotators across 450 head-to-head comparisons, ranked gpt-6-astra first with a 64.3% win rate, followed by grok-4.6 at 62.7% and muse-spark-1.3 at 60.3%. The Bradley-Terry model fit to the pairwise results returned the identical 1-to-10 order as the raw win-rate ranking, and seven different models finished first on at least one of the 10 design briefs. Kimi-k3 (53.3%) and glm-5.3 (53.2%) were treated as a tie, as were claude-sonnet-5 (49.1%) and qwen3.8-2.4t (48.9%), while seed-2.0-code finished last at 19.1%.", "body_md": "[Earlier this week](https://humansignal.com/blog/ai-has-a-house-style-we-measured-it-across-10-labs/), we gave 10 models from 10 labs the same 10 design briefs and prohibited the models from using references, frameworks, or web search. We wanted to understand the design sensibilities of the models themselves, including where they differ and converge, how they defend their design decisions, and how their designs align with human preferences.  We ended that post by saying people had also judged the designs. So here are those results: we put every design, every vote, and every stated rationale into an [interactive dashboard](https://humansignal-creative-services-report.netlify.app).\n\n  We ran a blind, balanced round-robin in [Label Studio](https://labelstud.io/).\n\nEvery model faced every other model once for each brief. That gave us 450 head-to-head comparisons. We collected 10 votes per comparison, for 4,500 votes from 29 annotators.\n\nAnnotators saw two sites side by side. They did not see the model names. They were tasked to pick which version they aesthetically preferred, or mark the round as a tie. Then, we calculated a global preference ranking using a win rate.\n\n  To calculate the win rate for model *i*, take every vote it was involved in. Give it a full point each time someone picked it, half a point each time someone called a tie, nothing when it lost. Divide the points by the number of votes. We show this as a formula below.\n\nThe following table shows the ranking of all ten models in the test set, ordered by win rate.\n\n| Model rankings |  |  |  |  | \n|---|---|---|---|---|\n| Rank | Model | Win rate | Bradley-Terry strength | Range across briefs | \n|---|---|---|---|---|\n| 1 | gpt-6-astra | 64.3% | 1.78 | 1st to 8th | \n| 2 | grok-4.6 | 62.7% | 1.67 | 1st to 8th | \n| 3 | muse-spark-1.3 | 60.3% | 1.52 | 1st to 7th | \n| 4 | deepseek-v4-pro | 56.2% | 1.29 | 1st to 9th | \n| 5 | kimi-k3 | 53.3% | 1.16 | 1st to 8th | \n| 6 | glm-5.3 | 53.2% | 1.15 | 1st to 10th | \n| 7 | claude-sonnet-5 | 49.1% | 0.98 | 4th to 9th | \n| 8 | qwen3.8-2.4t | 48.9% | 0.98 | 1st to 10th | \n| 9 | gemini-3.8-flash | 32.8% | 0.51 | 4th to 10th | \n| 10 | seed-2.0-code | 19.1% | 0.26 | 7th to 10th | \n\nTwo checks make us more confident in this ordering.\n\nFirst, we balanced the schedule. Every model faced the same opponents the same number of times. No model got an easier path through the study.\n\nSecond, we fit a Bradley-Terry model to the pairwise results. Bradley-Terry estimates for each model a single hidden strength, chosen so that the odds of one model beating another match their strength ratio across all 4,500 votes. We ran it because win rate ignores who a model faced: if our order were an artifact of scheduling, the two methods would disagree, and instead they return the identical 1-to-10 ranking. The Bradley-Terry scores return the same 1-to-10 order as our preference ranking, proving that our ranking isn’t the result of an unbalanced matchup or the scheduling, but rather is the result of true human preference.\n\nTwo pairs are too close to call. Kimi-k3 and glm-5.3 are separated by 0.1 percentage points. Claude-sonnet-5 and qwen3.8-2.4t are separated by 0.2.\n\nWe treat both pairs as ties.\n\nBefore you quote the overall ranking to claim which model did best, look at what happened on individual briefs.\n\nSeven different models finished first on at least one of the 10 briefs.\n\ngpt-6-astra won the overall study and won two briefs outright. It also finished 8th on the medical-research-conference brief.\n\nglm-5.3 had the best average brief rank of any model outside the top three, at 4.5. It still finished 6th overall because one poor result can outweigh several strong ones. On the science-publication brief, it finished last with a 24% win rate.\n\nclaude-sonnet-5 shows the other side of the distribution. It never finished above 4th or below 9th on any brief. That made it the most consistent model in the study.\n\nAnd… It still finished 7th. Consistency did not make it anyone's favorite.\n\nThis is why one number per model can be misleading. If you are choosing a model for a particular kind of design work, the overall ranking is less useful than the results for the briefs that look like your work.\n\nThe dashboard has four sections.\n\n**The designs**\n\nPick a brief from the thumbnail strip and browse the ten sites the models built for it, ranked by how often people chose each one. Open any design to see its cost, generation time, the model's stated direction, and the human evaluation. From there you can open the live site.\n\n**The decisions**\n\nThis is what the models decided before they wrote any code. For the same brief, compare the typefaces they named, the palettes by hue and saturation, and the reasoning they gave for those choices, in each model's own words.\n\n**Overall rankings**\n\nSee the global ranking across all ten briefs, including each model's best and worst finish, and how it compares with the pilot run.\n\n**Shared patterns**\n\nSee how far each model's choices sit from the rest, then where they converged. Inter appeared in 53 of 100 creative directions. IBM Plex Mono appeared in 34. Fraunces appeared in 25. The dashboard also shows the recurring color and layout choices.\n\nThe results are downloadable as JSON, and the rankings and pilot methodology as CSV.\n\nIf you want to check our arithmetic, the underlying numbers are there.\n\nIn the weeks to come, we’ll be deep diving into some of the most interesting results from this human preference study, backed by real data from the Dashboard. Stay tuned.", "url": "https://wpnews.pro/news/the-industry-trained-an-ai-design-monoculture-people-preferred-the-models-that", "canonical_source": "https://humansignal.com/blog/the-industry-trained-an-ai-design-monoculture-people-preferred-the-models-that-broke-from-it", "published_at": "2026-10-08 00:00:00+00:00", "updated_at": "2026-10-08 20:48:50.686725+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-research", "large-language-models"], "entities": ["HumanSignal", "Label Studio", "gpt-6-astra", "grok-4.6", "muse-spark-1.3", "deepseek-v4-pro", "kimi-k3", "glm-5.3"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-industry-trained-an-ai-design-monoculture-people-preferred-the-models-that", "markdown": "https://wpnews.pro/news/the-industry-trained-an-ai-design-monoculture-people-preferred-the-models-that.md", "text": "https://wpnews.pro/news/the-industry-trained-an-ai-design-monoculture-people-preferred-the-models-that.txt", "jsonld": "https://wpnews.pro/news/the-industry-trained-an-ai-design-monoculture-people-preferred-the-models-that.jsonld"}}