AI has a house style. We measured it across 10 labs. A study of 10 models from 10 labs across 10 design briefs found that AI-generated websites converge on a narrow house style, with telemedicine scoring 0.95 on primary-hue concentration (where 1.0 means every model chose the same hue) while independent film scored 0.14 and skatewear 0.08, a figure that masks 8 of 10 models choosing a near-black primary. HumanSignal ran the study using Flue for a common agent environment and OpenRouter for a pinned-provider interface, then collected 4,500 votes from 29 annotators across 450 blind, balanced round-robin matchups in Label Studio. HumanSignal CEO Michael Malyuk frames the finding as the emerging "taste moat": as AI makes execution cheaper, judgment about which version is worth making becomes the scarce advantage. Ask an AI model to design a modern website with little art direction and you can usually guess what you’re going to get. A soft gradient behind the hero. Rounded cards with drop shadows. Inter for the body copy. An expressive serif for the headline. You’ve seen it before. That raises a more interesting question than whether a model can build a website at all: what does it choose when nobody tells it what good looks like? Michael Malyuk, HumanSignal’s CEO, has been writing about this as the emerging “ taste moat https://www.linkedin.com/pulse/taste-moat-michael-malyuk-qrm3c/ ”: as AI makes execution cheaper and more accessible, judgment becomes more valuable. The hard part is no longer just getting something made. It’s knowing which version is worth making. We wanted to see what happens when you put that idea to the test. So we ran a study: 10 models from 10 labs, 10 design briefs, and a record of the creative decisions each model made along the way. Today’s models can build responsive, working sites in minutes. Most AI design benchmarks test whether they can reproduce something that already exists: give a model a reference interface and see how closely it can rebuild it. That tells us whether the model can execute. It doesn’t tell us much about its taste. So we took away the reference design. Each model had to make its own decisions about color, type, and layout, without the ability to use outside tools or packages. We also asked it to explain those decisions. Then we showed the results to people and asked which ones they preferred. This distinction matters beyond design. Once a model can reliably do something, the interesting question becomes what it chooses to do. Which explanation is clearer? Which slide works better? When should a voice assistant interrupt? Which of two perfectly functional interfaces would you actually want to use? As Michael put it in his writing on taste, the advantage increasingly shifts from simply having access to capable models to knowing what good looks like. That’s a much harder thing to evaluate. There’s no compiler for taste. We gave 10 models from 10 labs the same 10 briefs, ranging from a telemedicine app to a New York skatewear brand. We used Flue to give every model the same agent environment and OpenRouter for a common interface, pinning each request to a named provider with fallbacks disabled. For each brief, every model went through three arms: All 10 models used the same image-generation model, so differences in image quality wouldn’t explain the results. Then we asked people to judge the results. We ran a blind, balanced round-robin in Label Studio. Every pair of models went head-to-head on every brief, giving us 450 matchups and 4,500 votes from 29 annotators. Annotators saw two sites side by side. They didn’t know which model made either one. They picked a winner or called it a tie. The goal was simple: separate “this works” from “I prefer this.” That distinction is where taste becomes measurable. Across 100 generated sites, three patterns show up again and again: Color showed a similar pattern. For the telemedicine brief, the models scored 0.95 on primary-hue concentration , where 1.0 means every model chose the same hue. Independent film was the opposite. It scored 0.14, making it the most varied brief in the study. Skatewear scored just 0.08. That number sounds like maximum disagreement, but it hides one of the strongest points of agreement in the entire study: 8 of 10 models chose a near-black primary and largely stayed away from color altogether. When models did disagree, they tended to split into a couple of predictable camps. Boutique hotel split between warm gold and deep green. Dog park bar split between orange and green. Independent film split between amber and blue. Ten different labs, and we keep ending up with the same two or three answers. FIGURE: palette swatch matrix from findings-pilot-001.html The most interesting part is that this convergence doesn’t begin with CSS. In the direction arm, there was no code, framework, or component library. The models were simply describing what they wanted to make. The convergence was already there. Two of the 10 models emitted Tailwind color tokens. The other eight didn’t. But all 10 still landed in roughly the same place. Their reasoning looked similar, too. Of the 563 color rationales, 36% described a choice in terms of what the model wanted to avoid: “without X,” “rather than Y,” and similar constructions. The models were trying to differentiate their designs. And they kept arriving at similar ideas of what “different” should look like. That’s the part of the taste-moat argument that we wanted to make concrete. If every model has access to roughly the same capabilities, differentiation has to come from somewhere else. And if the models’ default judgments converge, then capability alone isn’t going to produce distinctive work. We’ll dig into those rationales—and the references behind them—in a later post. If your product generates design, this house style is likely part of the default. And a traditional competence benchmark probably won’t tell you that. These sites work. They’re responsive. They’re polished. They pass. They can also look remarkably similar. That creates a problem for anyone building products around generative AI: you can measure whether the model can do the job without measuring whether people actually prefer what it did. That’s the gap we wanted to study. If taste is becoming a moat, then taste needs an evaluation loop. Show people the outputs. Ask them to choose. Measure where they agree, where they disagree, and which models consistently make choices people prefer. That’s what we did here. Next, we’ll publish what our annotators chose across every blind matchup—and which models came out ahead. Stay tuned.