{"slug": "what-your-design-system-can-t-teach-ai-agents", "title": "What your design system can't teach AI agents", "summary": "Design systems cannot teach AI agents composition, the judgment of what goes on a screen and what gets left out, because that judgment is applied by hand and passed on through proximity rather than documentation, according to Hamza Oza. Oza, who works with the agent-driven software factory Tessl and Kikimora, codified brand values and voice into checkable rule documents for agents but found the visual pillar hardest to write down, as agent-built screens passed every linter while still stacking panels with four levels of elevation and interrupting users with modals. Oza extracted design rules by comparing three agent-built screens in a Figma diff after introspection failed.", "body_md": "ARTICLE\n\n# What your design system can't teach AI agents\n\nHamza Oza\n\nFor months, colleagues kept asking me a version of the same question. Could I write down how I approach design, so they could draw on it without waiting on my calendar?\n\nThe question kept coming back because of how we now build. [Tessl runs a software factory, Kikimora](https://www.youtube.com/watch?v=u37qkpp5eB8): agents do the making, people set direction and decide what ships. As its use spread, more of what we shipped was arriving without a designer having touched it.\n\nWhat people ask for is almost always help with the pixels, but you cannot start there because a brand comprises three pillars:\n\n- **Values** are what an organization believes.\n- **Voice** is how it articulates those beliefs.\n- **Visuals** are how it presents them.\n\nCodify the visuals alone and an agent gets a list of preferences with none of the reasoning underneath. The order is forced, even though the pressure comes entirely from the visual end. It also reinforces that design only equates to pixels.\n\nVisuals are also the hardest to write down, and the reason took me a while to name. Composition, the judgment of what goes on a screen, where it sits, and what gets left out, lives in a designer's head. It is applied by hand and passed on through proximity, not documentation, because a person was always there to supply it. Nobody writes down what nobody has ever had to hand over.\n\n## Codifying values and voice for agents\n\nVoice was the most tractable. Every piece of copy the organization has ever shipped is a data point, so the rules are recoverable: 'American English throughout', 'name the thing before explaining it', 'no exclamation marks', 'no em dashes'. The first pass at the rules came straight from that shipped copy, then each one was refined with our Head of Marketing through worked examples until it met our standards. A person can check these. So can a machine.\n\nValues look easier but are not. Most organizations have written theirs down already, but the words are open to many interpretations. The real work is spelling out the decisions each value implies, which is where everyone discovers they were not agreeing after all.\n\nValues started from our operating principles, and I worked through each one with our operations and leadership teams, pushing until it produced a decision without losing the spirit of the principle it came from. The output of both is deliberately plain: a document of checkable rules, which is exactly the form an agent can act on.\n\n## Why a design system can't teach composition\n\nWhich is usually where someone asks: but what about the design system?\n\nWe have one. It is good. But it was never going to teach composition. A component library tells you what the parts are. It says nothing about which parts to reach for, how many, or when the right answer is none of them. Composition is its own discipline, and no designer expects the library to carry it.\n\nAgents make that gap expensive. They inherit the library in full and the judgment not at all, and they will not absorb it by sitting near me for six months.\n\nThis was exemplified by the screens our software factory was producing. Every component came from our design system, every color was a token, every spacing value was legal. Nothing would have failed a linter. Yet panels sat on panels with four levels of elevation, each drop shadow implying a surface floating above the last, and no hierarchy telling the eye where to start. Borders wherever two things met, as though adjacency needed explaining. Modals interrupting people for no good reason.\n\nOn the surface, none of these are component bugs. Each is a composition decision, a call I make in seconds without conscious effort. Judgment that fast never produces an artifact, so there was nothing for an agent to read.\n\n## Extracting design rules from a Figma diff\n\nHaving failed at introspection, I went the other way round. I took three agent-built screens, deliberately different in purpose and layout, and redesigned each properly in Figma. Then I put each pair in front of another agent and asked it to describe the difference: not to judge which was better, but to enumerate what had changed.\n\nLow expectations. Mostly I wanted to see what it would come up with.\n\nEach pair came back as a list of specific, checkable deltas. Nested surfaces, cards sitting inside cards, cut from three levels to one. Fourteen borders removed, separation carried by spacing. A modal replaced by an inline expansion. Section spacing doubled. Most were decisions I had made in one fell swoop, and seeing them itemized was the first time my own reasoning had been visible to me.\n\nThe process works best across several designs, because that is how you tell a rule from a one-off. A change recurring across all three was an emerging rule.\n\nDiffs suit how agents work. Introspection asks you to retrieve something that was never stored as language. A diff turns it into an observable difference between two artifacts, which is exactly what a model is good at describing.\n\nNone of this makes design a checklist. Good design is contextual, and part of the craft is knowing when to break your own rule. When these rules later become checks, they should start at 'warn' rather than 'block': in a discipline where the exceptions carry information, blocking the exception throws away your most interesting signal.\n\n## What automated design and brand checks actually catch\n\nGuidance sitting in a document is not guidance that acts. Ours ships as two skills in a plugin: one loads at the point of creation, before the agent makes anything, and the other reviews output against the same rules at the pull request and again before release.\n\nThe trigger is nothing cleverer than the skill's own description: when a task involves making something user-facing, a screen, a doc, a piece of copy, the agent matches on that and loads the skill before it writes a line. Each skill pulls in the pillar docs the task touches, resolving conflicts in a fixed order: values decide ambiguous calls, voice governs the copy, visuals govern layout and hierarchy. Both end with a checklist, and an instruction to surface any call that needs human judgment rather than guess at it.\n\nWhat I did not anticipate was pointing the same guidance at work that already existed. We built automations to sweep whole surfaces rather than check new changes, and the first run raised thirteen pull requests against accumulated drift. All thirteen merged.\n\nSome were mundane. Our external copy standard is American English, but the team is mostly London-based and writes British by reflex, so 'colour' and 'organisation' keep turning up in the product. A reviewer who spells it 'colour' will not flag it; the agent has no such blind spot. Other catches were harder, where the copy was grammatically fine but did not sound like us.\n\nNone of the thirteen would have been caught by a reviewer reading a diff, because none arrived in one. They accumulated a word at a time, each change too small to object to. A sweep sees what a diff-scoped check structurally cannot. It now runs weekly across our codebases.\n\n## Is it complete? Of course not.\n\nThe rules are incomplete and some will turn out to be wrong. That is not a temporary state I am working through. Maintaining the guidance, and the agents that enforce it, is now part of the discipline: the rules change as the product changes, the sweeps surface things that need new rules, and none of it arrives at done.\n\nIt would be easy to read all this as a designer writing themselves out of the work. It is closer to the opposite. An agent can apply a rule to every screen the factory produces, but it cannot decide what the rule should be, or whether the exception in front of it is a mistake or the interesting kind. Deciding the guidelines and verifying what comes back is where a designer's judgment now concentrates, and it is leverage of a sort the job never had: one person's taste, reaching every screen that ships.\n\nIf you want to try this yourself, the path does not start with a manifesto. Take three screens an agent built, redesign them by hand, and ask a model to enumerate the differences; what recurs across all three is your first rule set. Ship those rules to the point of creation, set the gates to warn, and point a sweep at something that already exists.\n\nIt is better than unguided agent output, and it improves each time a correction goes into the guidance rather than a single artifact. That is the test I hold it to: not whether the line produces perfect work, but whether it produces better work this month than last.\n\nWhich is not a new job, in the end. I trained as a mechanical engineer, and factories run on one principle: quality is built into the line, not inspected at the end of it. The craft moving from the artifact to what the artifact is made from is that same idea, turning up where I did not expect it.\n\nCOPY & SHARE\n\nHamza Oza\n\nHamza Oza is a Designer at Tessl building tools for complex technical domains and a Visiting Tutor at the Royal College of Art working at the intersection of design and technology.\n\nREADING\n\n·\n\n0%\n\nCOPY & SHARE\n\nHamza Oza\n\nHamza Oza is a Designer at Tessl building tools for complex technical domains and a Visiting Tutor at the Royal College of Art working at the intersection of design and technology.\n\nYOUR NEXT READ\n\n## Reflection Before Augmentation\n\nExploring the challenges of integrating AI into teams, emphasizing the importance of reflection and context engineering before adopting new tools for organizational success.\n\nHamza Oza", "url": "https://wpnews.pro/news/what-your-design-system-can-t-teach-ai-agents", "canonical_source": "https://tessl.io/blog/what-your-design-system-cant-teach-ai-agents", "published_at": "2026-09-02 12:00:00+00:00", "updated_at": "2026-09-23 20:31:21.030504+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "generative-ai"], "entities": ["Hamza Oza", "Tessl", "Kikimora", "Figma"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-your-design-system-can-t-teach-ai-agents", "markdown": "https://wpnews.pro/news/what-your-design-system-can-t-teach-ai-agents.md", "text": "https://wpnews.pro/news/what-your-design-system-can-t-teach-ai-agents.txt", "jsonld": "https://wpnews.pro/news/what-your-design-system-can-t-teach-ai-agents.jsonld"}}