{"slug": "we-stopped-letting-the-model-write-jsx", "title": "We Stopped Letting the Model Write JSX", "summary": "A developer replaced an AI support case composer that let a language model stream React/JSX snippets into a transpiler and new Function with a declarative A2UI v0.9 architecture, in which the model emits JSON UI descriptions against a host-owned component catalog instead of executable code. The v1 approach produced inconsistent styling, hallucinated components, malformed markdown tables and an unacceptable prompt-injection-to-code-execution trust boundary; the new design separates surface creation, component descriptions and data-model updates, with no onClick, style or dangerouslySetInnerHTML fields available to the model.", "body_md": "*How a support case composer went from \"suggest a React snippet\" to a declarative catalog with A2UI v0.9, and what broke along the way*\n\nThe scariest pull request I reviewed last quarter was four lines long.\n\nIt took a string the model had streamed back, ran it through a transpiler, and handed the result to new Function. The author (a good engineer, on a deadline) had a comment above it: // sandboxed enough for internal tools. It wasn't. I left a review comment I'm still a little proud of, closed the laptop, and went for a walk.\n\nThat PR was the end of v1 of our support case composer. This is the story of what we replaced it with.\n\nv1: let the model cook\n\nThe feature sounded simple. A support agent opens a case, clicks \"Compose\", and an AI helper builds a working view for that case: customer summary, recent timeline, a draft reply, maybe a small table of related tickets.\n\nOur first version asked the model for markdown, plus the occasional fenced React snippet when it wanted something richer than a table. For about two weeks it felt like magic. Then staging started teaching us things.\n\nThe model had no idea what our design system was. It would write a perfectly reasonable with Tailwind classes we don't use. It invented a that didn't exist. The same prompt on Monday and Thursday produced two different visual dialects.\n\nMarkdown tables are not a UI. A row would arrive with six cells in a five-column table. A \"status\" column would be an emoji one time and a word the next. There was nothing to bind an action to, so every interactive thing the model wanted turned into the sentence \"you could click here.\"\n\nAnd then there was the trust boundary. Customer messages are untrusted input. Customer messages flow into the model's context. The model's output flows into our renderer. Draw that line on a whiteboard and you'll feel your stomach drop. We never shipped the eval, but we were one sprint away from it, and that's the part that bothered me. Prompt injection plus code execution isn't a bug you patch. It's an architecture you refuse.\n\nThe shift: the model emits data, the host owns the pixels\n\nWe moved to A2UI v0.9. The idea, in one sentence: the agent doesn't send UI code, it sends a description of a UI as JSON, and our app decides how (and whether) to render each piece.\n\nThe wire format is a stream of small messages. The three we use constantly:\n\n```\n{ \"version\": \"v0.9\", \"createSurface\": { \"surfaceId\": \"case-4821\", \"catalogId\": \"support.internal:case-composer\" } }\n{\n  \"version\": \"v0.9\",\n  \"updateComponents\": {\n    \"surfaceId\": \"case-4821\",\n    \"components\": [\n      { \"id\": \"root\", \"component\": \"Column\", \"children\": [\"summary\", \"reply\"] },\n      { \"id\": \"summary\", \"component\": \"CaseSummary\", \"caseId\": \"4821\" },\n      { \"id\": \"reply\", \"component\": \"DraftReply\", \"body\": { \"path\": \"/draft/body\" } }\n    ]\n  }\n}\n{ \"version\": \"v0.9\", \"updateDataModel\": { \"surfaceId\": \"case-4821\", \"path\": \"/draft\", \"value\": { \"body\": \"Hi Dana, thanks for your patience...\" } } }\n```\n\nA surface gets created, components get described, and data gets poured in separately. That separation turned out to matter more than I expected. The structure of the view and the content of the view are different messages, so the model can fix a typo in the draft without re-describing the whole layout.\n\nNotice what's missing: there's no onClick, no style, no dangerouslySetInnerHTML. There's nowhere to put them.\n\nThe catalog is the contract\n\nThe catalogId in createSurface points at something we own: a catalog of the components the agent is allowed to ask for. If it's not in the catalog, it doesn't exist as far as the model is concerned.\n\nWe define each component as a name plus a Zod schema, then bind it to a real React renderer. The exact shape of the binding depends on your @a2ui/react version, so check the docs for the current signature. The gist looks like this:\n\n``` python\nimport { z } from 'zod';\nimport { Catalog } from '@a2ui/web_core/v0_9';\nimport type { ReactComponentImplementation } from '@a2ui/react/v0_9';\n\n// The contract: what the agent may say about a CaseSummary\nexport const CaseSummarySchema = z.object({\n  caseId: z.string().regex(/^\\d+$/),\n  showTimeline: z.boolean().optional(),\n});\n\n// The renderer: what *we* decide it looks like\nfunction CaseSummary({ caseId, showTimeline }: z.infer<typeof CaseSummarySchema>) {\n  const data = useCase(caseId); // our own data layer, our own auth\n  return <SummaryCard data={data} timeline={showTimeline} />;\n}\n\nexport const caseComposerCatalog = new Catalog<ReactComponentImplementation>(\n  'support.internal:case-composer', // must match createSurface.catalogId\n  [/* CaseSummary, DraftReply, Timeline, RelatedTickets, ... bound to their schemas */],\n);\n```\n\nTwo things I like about this.\n\nFirst, the schema is both validation and documentation. We hand the same definitions to the model as part of its prompt, so what the agent is told and what the host enforces can't drift apart.\n\nSecond, the renderer fetches its own data. The agent says CaseSummary caseId=\"4821\". It never carries the customer's email address in the message, and useCase runs through our normal permission checks. The model can request a component, but it can't widen what the viewer is allowed to see.\n\nWiring up the MessageProcessor\n\nOn the client, the plumbing is small. The MessageProcessor from @a2ui/web_core takes the catalogs you support and turns the raw message stream into surface state that the React layer renders:\n\n``` python\nimport { MessageProcessor } from '@a2ui/web_core/v0_9';\nimport type { ReactComponentImplementation } from '@a2ui/react/v0_9';\n\nconst processor = new MessageProcessor<ReactComponentImplementation>([caseComposerCatalog]);\n\nasync function consume(stream: AsyncIterable<unknown>) {\n  for await (const message of stream) {\n    const verdict = guard(message);          // our check, more on this below\n    if (!verdict.ok) { report(verdict); continue; }\n    processor.processMessages([message]);    // check your version's exact method name\n  }\n}\n```\n\nThe agent talks to us over server-sent events, one JSON message per event. We feed each one through our guard and then into the processor. The @a2ui/react surface subscribes to the processor's state and re-renders as pieces land. Nothing clever, which is exactly what I wanted at the boundary.\n\nRefusing unknown components\n\nThe catalog already limits what can render. But \"can't render\" and \"fails loudly and safely\" are different properties, and the second one is where we spent a staging week.\n\nWe put our own guard in front of the processor. It's host code, not a package feature, and it does three boring things: checks the catalogId is one we registered, checks every component name against the allowlist, and runs each component's props through its Zod schema.\n\n``` js\nconst allowed = new Set(Object.keys(componentSchemas));\n\nfunction guard(msg: any): { ok: true } | { ok: false; reason: string; detail?: unknown } {\n  if (msg.createSurface && msg.createSurface.catalogId !== CATALOG_ID) {\n    return { ok: false, reason: 'unknown-catalog' };\n  }\n  for (const c of msg.updateComponents?.components ?? []) {\n    if (!allowed.has(c.component)) {\n      return { ok: false, reason: 'unknown-component', detail: c.component };\n    }\n    const { component, id, children, ...props } = c;\n    const parsed = componentSchemas[component].safeParse(props);\n    if (!parsed.success) {\n      return { ok: false, reason: 'invalid-props', detail: parsed.error.flatten() };\n    }\n  }\n  return { ok: true };\n}\n```\n\nWhat broke in staging: in the first week, the model asked for a Chart component about once every twenty surfaces. We didn't have one. It's a reasonable thing to want in a case view, so the model kept reaching for it.\n\nWith the guard, that message gets dropped and logged with the component name. The surface still renders everything else, plus a small neutral \"couldn't display one element\" placeholder. No crash, no blank screen, no mystery.\n\nThe log turned out to be a product-roadmap tool. We literally sorted unknown-component by frequency and built the top one. (It was Chart. Of course it was.)\n\nOne more lesson here: fail the element, not the surface. My first version rejected the whole message on any bad component, and a single malformed prop blanked the entire view. Dropping just the offending piece, and keeping the rest, was much kinder to the people actually using it.\n\nStreaming surfaces\n\nBecause updateComponents is a flat list of components that reference each other by id, a parent can arrive before the children it points at. That's great for perceived speed. The summary card shows up while the model is still thinking about the draft reply.\n\nIt also gave us our most embarrassing bug. A Column listing [\"summary\", \"reply\"] would render with a hole where reply hadn't arrived yet, the layout would jump when it did, and on a slow connection the whole thing looked like it was having a small breakdown.\n\nThe fix was dull and effective: a skeleton for any child id that doesn't resolve yet, with a fixed minimum height so nothing shifts. Streaming went from \"flickery demo\" to \"feels fast\", and it cost us about twenty lines.\n\nThe button that wanted to send an email\n\nThe moment I stopped worrying and started trusting the design came from something the agent did on its own.\n\nA few days after launch, a support lead pinged me with a screenshot. The agent had put a \"Send reply to customer\" button into a surface. We hadn't asked for that. It's a sensible thing for a composer to want, but it's a side effect: an email, to a real customer, from a real person's account.\n\nWe hadn't planned for it, but the catalog design meant the answer was already structural. Our ActionButton component only accepts an intent from a fixed enum, and the host decides what each intent is allowed to do. Intents are split into two classes:\n\nRead-only (open a ticket, expand a timeline, copy text) run immediately.\n\nSide-effecting (send reply, close case, issue refund) never run from a click on an agent-authored button. They open a confirmation panel showing exactly what will happen, with the final text and the recipient, and a human has to approve it.\n\nAnd the server re-checks. The click produces a request that our backend validates against the agent's permissions and the person's role, regardless of what the UI claimed. The button was a proposal. The approval was the action.\n\nI want to be clear that none of this is \"AI safety\" in the fancy sense. It's the same discipline you'd apply to any untrusted client: the front end suggests, the back end decides. The difference is that the declarative catalog gave us a single place to enforce it.\n\nWhat we'd do differently\n\nWrite the guard before the first prompt. We built it after the incident in staging. It should have existed on day one.\n\nVersion the catalog. We changed a prop name and old in-flight surfaces failed validation. Put a version in the catalogId and keep the previous one registered for a while.\n\nStart with fewer components. We began with fourteen. Six did ninety percent of the work, and the other eight made the prompt longer and the model more creative in ways we didn't want.\n\nLog every refusal with the full message. Not just the component name. When something looks weird in a screenshot, you want the receipts.\n\nWhen this is worth it, and when it isn't\n\nDeclarative generative UI is not free. You're building a catalog, a validation layer, a confirmation pattern, and a prompt that teaches the model your vocabulary. For some products that's overkill.\n\nUse a fixed tool-rendered component map when the set of views is small and known. If the agent calls get_case and you always render a CaseCard, you don't need a protocol. Map tool names to components, pass typed results, and go home early. That's simpler, easier to test, and perfectly good.\n\nReach for generative UI when the composition varies and you can't predict it: the right view depends on the case, the person, and what the agent found. That's where \"the agent picks and arranges from an approved set\" earns its keep, and where hand-writing a component for every combination collapses under its own weight.\n\nMy recommendation: start with the fixed map. Ship it. Watch where users ask for variations the map can't express. When you see that pattern repeating, introduce a catalog for those surfaces only, behind the same guard and the same human-approval rule for side effects. Don't let the model write code, and don't let it near anything that mutates the world without a person in the loop.\n\nThe best thing about our v2 isn't the UI the model can produce. It's the UI it can't.", "url": "https://wpnews.pro/news/we-stopped-letting-the-model-write-jsx", "canonical_source": "https://dev.to/faisalinfinity/we-stopped-letting-the-model-write-jsx-fik", "published_at": "2026-10-02 19:18:24+00:00", "updated_at": "2026-10-02 19:37:52.087755+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "generative-ai", "ai-tools", "developer-tools"], "entities": ["A2UI", "React", "Zod", "Tailwind"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-stopped-letting-the-model-write-jsx", "markdown": "https://wpnews.pro/news/we-stopped-letting-the-model-write-jsx.md", "text": "https://wpnews.pro/news/we-stopped-letting-the-model-write-jsx.txt", "jsonld": "https://wpnews.pro/news/we-stopped-letting-the-model-write-jsx.jsonld"}}