cd /news/ai-agents/we-stopped-letting-the-model-write-j… · home › topics › ai-agents › article
[ARTICLE · art-144099] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

We Stopped Letting the Model Write JSX

A developer replaced an AI support case composer that let a language model stream React/JSX snippets into a transpiler and new Function with a declarative A2UI v0.9 architecture, in which the model emits JSON UI descriptions against a host-owned component catalog instead of executable code. The v1 approach produced inconsistent styling, hallucinated components, malformed markdown tables and an unacceptable prompt-injection-to-code-execution trust boundary; the new design separates surface creation, component descriptions and data-model updates, with no onClick, style or dangerouslySetInnerHTML fields available to the model.

by read10 min views1 publishedOct 2, 2026

How a support case composer went from "suggest a React snippet" to a declarative catalog with A2UI v0.9, and what broke along the way

The scariest pull request I reviewed last quarter was four lines long.

It took a string the model had streamed back, ran it through a transpiler, and handed the result to new Function. The author (a good engineer, on a deadline) had a comment above it: // sandboxed enough for internal tools. It wasn't. I left a review comment I'm still a little proud of, closed the laptop, and went for a walk.

That PR was the end of v1 of our support case composer. This is the story of what we replaced it with.

v1: let the model cook

The feature sounded simple. A support agent opens a case, clicks "Compose", and an AI helper builds a working view for that case: customer summary, recent timeline, a draft reply, maybe a small table of related tickets.

Our first version asked the model for markdown, plus the occasional fenced React snippet when it wanted something richer than a table. For about two weeks it felt like magic. Then staging started teaching us things.

The model had no idea what our design system was. It would write a perfectly reasonable with Tailwind classes we don't use. It invented a that didn't exist. The same prompt on Monday and Thursday produced two different visual dialects.

Markdown tables are not a UI. A row would arrive with six cells in a five-column table. A "status" column would be an emoji one time and a word the next. There was nothing to bind an action to, so every interactive thing the model wanted turned into the sentence "you could click here."

And then there was the trust boundary. Customer messages are untrusted input. Customer messages flow into the model's context. The model's output flows into our renderer. Draw that line on a whiteboard and you'll feel your stomach drop. We never shipped the eval, but we were one sprint away from it, and that's the part that bothered me. Prompt injection plus code execution isn't a bug you patch. It's an architecture you refuse.

The shift: the model emits data, the host owns the pixels

We moved to A2UI v0.9. The idea, in one sentence: the agent doesn't send UI code, it sends a description of a UI as JSON, and our app decides how (and whether) to render each piece.

The wire format is a stream of small messages. The three we use constantly:

{ "version": "v0.9", "createSurface": { "surfaceId": "case-4821", "catalogId": "support.internal:case-composer" } }
{
  "version": "v0.9",
  "updateComponents": {
    "surfaceId": "case-4821",
    "components": [
      { "id": "root", "component": "Column", "children": ["summary", "reply"] },
      { "id": "summary", "component": "CaseSummary", "caseId": "4821" },
      { "id": "reply", "component": "DraftReply", "body": { "path": "/draft/body" } }
    ]
  }
}
{ "version": "v0.9", "updateDataModel": { "surfaceId": "case-4821", "path": "/draft", "value": { "body": "Hi Dana, thanks for your patience..." } } }

A surface gets created, components get described, and data gets poured in separately. That separation turned out to matter more than I expected. The structure of the view and the content of the view are different messages, so the model can fix a typo in the draft without re-describing the whole layout.

Notice what's missing: there's no onClick, no style, no dangerouslySetInnerHTML. There's nowhere to put them.

The catalog is the contract

The catalogId in createSurface points at something we own: a catalog of the components the agent is allowed to ask for. If it's not in the catalog, it doesn't exist as far as the model is concerned.

We define each component as a name plus a Zod schema, then bind it to a real React renderer. The exact shape of the binding depends on your @a2ui/react version, so check the docs for the current signature. The gist looks like this:

import { z } from 'zod';
import { Catalog } from '@a2ui/web_core/v0_9';
import type { ReactComponentImplementation } from '@a2ui/react/v0_9';

// The contract: what the agent may say about a CaseSummary
export const CaseSummarySchema = z.object({
  caseId: z.string().regex(/^\d+$/),
  showTimeline: z.boolean().optional(),
});

// The renderer: what *we* decide it looks like
function CaseSummary({ caseId, showTimeline }: z.infer<typeof CaseSummarySchema>) {
  const data = useCase(caseId); // our own data layer, our own auth
  return <SummaryCard data={data} timeline={showTimeline} />;
}

export const caseComposerCatalog = new Catalog<ReactComponentImplementation>(
  'support.internal:case-composer', // must match createSurface.catalogId
  [/* CaseSummary, DraftReply, Timeline, RelatedTickets, ... bound to their schemas */],
);

Two things I like about this.

First, the schema is both validation and documentation. We hand the same definitions to the model as part of its prompt, so what the agent is told and what the host enforces can't drift apart.

Second, the renderer fetches its own data. The agent says CaseSummary caseId="4821". It never carries the customer's email address in the message, and useCase runs through our normal permission checks. The model can request a component, but it can't widen what the viewer is allowed to see.

Wiring up the MessageProcessor

On the client, the plumbing is small. The MessageProcessor from @a2ui/web_core takes the catalogs you support and turns the raw message stream into surface state that the React layer renders:

import { MessageProcessor } from '@a2ui/web_core/v0_9';
import type { ReactComponentImplementation } from '@a2ui/react/v0_9';

const processor = new MessageProcessor<ReactComponentImplementation>([caseComposerCatalog]);

async function consume(stream: AsyncIterable<unknown>) {
  for await (const message of stream) {
    const verdict = guard(message);          // our check, more on this below
    if (!verdict.ok) { report(verdict); continue; }
    processor.processMessages([message]);    // check your version's exact method name
  }
}

The agent talks to us over server-sent events, one JSON message per event. We feed each one through our guard and then into the processor. The @a2ui/react surface subscribes to the processor's state and re-renders as pieces land. Nothing clever, which is exactly what I wanted at the boundary.

Refusing unknown components

The catalog already limits what can render. But "can't render" and "fails loudly and safely" are different properties, and the second one is where we spent a staging week.

We put our own guard in front of the processor. It's host code, not a package feature, and it does three boring things: checks the catalogId is one we registered, checks every component name against the allowlist, and runs each component's props through its Zod schema.

const allowed = new Set(Object.keys(componentSchemas));

function guard(msg: any): { ok: true } | { ok: false; reason: string; detail?: unknown } {
  if (msg.createSurface && msg.createSurface.catalogId !== CATALOG_ID) {
    return { ok: false, reason: 'unknown-catalog' };
  }
  for (const c of msg.updateComponents?.components ?? []) {
    if (!allowed.has(c.component)) {
      return { ok: false, reason: 'unknown-component', detail: c.component };
    }
    const { component, id, children, ...props } = c;
    const parsed = componentSchemas[component].safeParse(props);
    if (!parsed.success) {
      return { ok: false, reason: 'invalid-props', detail: parsed.error.flatten() };
    }
  }
  return { ok: true };
}

What broke in staging: in the first week, the model asked for a Chart component about once every twenty surfaces. We didn't have one. It's a reasonable thing to want in a case view, so the model kept reaching for it.

With the guard, that message gets dropped and logged with the component name. The surface still renders everything else, plus a small neutral "couldn't display one element" placeholder. No crash, no blank screen, no mystery.

The log turned out to be a product-roadmap tool. We literally sorted unknown-component by frequency and built the top one. (It was Chart. Of course it was.)

One more lesson here: fail the element, not the surface. My first version rejected the whole message on any bad component, and a single malformed prop blanked the entire view. Dropping just the offending piece, and keeping the rest, was much kinder to the people actually using it.

Streaming surfaces

Because updateComponents is a flat list of components that reference each other by id, a parent can arrive before the children it points at. That's great for perceived speed. The summary card shows up while the model is still thinking about the draft reply.

It also gave us our most embarrassing bug. A Column listing ["summary", "reply"] would render with a hole where reply hadn't arrived yet, the layout would jump when it did, and on a slow connection the whole thing looked like it was having a small breakdown.

The fix was dull and effective: a skeleton for any child id that doesn't resolve yet, with a fixed minimum height so nothing shifts. Streaming went from "flickery demo" to "feels fast", and it cost us about twenty lines.

The button that wanted to send an email

The moment I stopped worrying and started trusting the design came from something the agent did on its own.

A few days after launch, a support lead pinged me with a screenshot. The agent had put a "Send reply to customer" button into a surface. We hadn't asked for that. It's a sensible thing for a composer to want, but it's a side effect: an email, to a real customer, from a real person's account.

We hadn't planned for it, but the catalog design meant the answer was already structural. Our ActionButton component only accepts an intent from a fixed enum, and the host decides what each intent is allowed to do. Intents are split into two classes:

Read-only (open a ticket, expand a timeline, copy text) run immediately.

Side-effecting (send reply, close case, issue refund) never run from a click on an agent-authored button. They open a confirmation panel showing exactly what will happen, with the final text and the recipient, and a human has to approve it.

And the server re-checks. The click produces a request that our backend validates against the agent's permissions and the person's role, regardless of what the UI claimed. The button was a proposal. The approval was the action.

I want to be clear that none of this is "AI safety" in the fancy sense. It's the same discipline you'd apply to any untrusted client: the front end suggests, the back end decides. The difference is that the declarative catalog gave us a single place to enforce it.

What we'd do differently

Write the guard before the first prompt. We built it after the incident in staging. It should have existed on day one.

Version the catalog. We changed a prop name and old in-flight surfaces failed validation. Put a version in the catalogId and keep the previous one registered for a while.

Start with fewer components. We began with fourteen. Six did ninety percent of the work, and the other eight made the prompt longer and the model more creative in ways we didn't want.

Log every refusal with the full message. Not just the component name. When something looks weird in a screenshot, you want the receipts.

When this is worth it, and when it isn't

Declarative generative UI is not free. You're building a catalog, a validation layer, a confirmation pattern, and a prompt that teaches the model your vocabulary. For some products that's overkill.

Use a fixed tool-rendered component map when the set of views is small and known. If the agent calls get_case and you always render a CaseCard, you don't need a protocol. Map tool names to components, pass typed results, and go home early. That's simpler, easier to test, and perfectly good.

Reach for generative UI when the composition varies and you can't predict it: the right view depends on the case, the person, and what the agent found. That's where "the agent picks and arranges from an approved set" earns its keep, and where hand-writing a component for every combination collapses under its own weight.

My recommendation: start with the fixed map. Ship it. Watch where users ask for variations the map can't express. When you see that pattern repeating, introduce a catalog for those surfaces only, behind the same guard and the same human-approval rule for side effects. Don't let the model write code, and don't let it near anything that mutates the world without a person in the loop.

The best thing about our v2 isn't the UI the model can produce. It's the UI it can't.

── more in #ai-agents 4 stories · sorted by recency
── more on @a2ui 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/we-stopped-letting-t…] indexed:0 read:10min 2026-10-02 · —