{"slug": "we-reverse-engineered-chatgpt-intelligent-ui", "title": "We reverse engineered ChatGPT Intelligent UI", "summary": "A technical analysis published in October 2026 details how ChatGPT's Intelligent UI, introduced with GPT-6, renders interactive components inline in conversations by having the model write interfaces in a language OpenAI calls DIL, which combines Markdown, JSX-like tags, and JavaScript. The backend compiles each partial response into a JavaScript program and a JSON document stored as model_dil_v2, wrapping expressions in __dilSafe for error isolation and assigning stable state keys so values survive recompilation, while the client executes the program in a sandbox and applies the resulting operations to ChatGPT's native components. The format is designed to be writable token by token and usable while half-written, letting the server cut a partial response at its last complete construct and still compile it.", "body_md": "## [What is Intelligent UI?](#what-is-intelligent-ui)\n\nIntelligent UI is the capability, introduced with GPT-6, that allows ChatGPT to respond with interactive interfaces rendered inline in the conversation ([OpenAI announcement](https://openai.com/index/gpt-6-for-everyone/)). A single response can combine prose with components such as sliders, forms, tables, charts, maps, and product cards. These components respond to input without a further model call, and they are drawn with ChatGPT's own design system rather than embedded as a web page.\n\nThe observations in this article were made on ChatGPT for the web in October 2026, using GPT-6 and GPT-6 Thinking.\n\n## [Building blocks](#building-blocks)\n\nChatGPT's implementation divides the work between the model, the backend server, and the client:\n\n- [Inference format](#inference-format) : the model writes the interface in DIL, which combines Markdown with JSX-like tags and JavaScript.\n- [Server-side compilation](#server-side-compilation) : the server converts each partial response into a JavaScript program and a JSON document of text and data.\n- [Client runtime](#client-runtime) : a sandboxed runtime executes the program and produces UI operations.\n- [Rendering](#rendering) : ChatGPT applies the operations to its own native components.\n- [Design system and catalog](#design-system-and-catalog) : the components, properties, and design tokens available to the model.\n\n## [Inference format](#inference-format)\n\nThis is what the model writes. In ChatGPT, it is a language OpenAI calls DIL: Markdown for prose, JSX-like tags for components, and JavaScript for state and logic. We will follow one small response through every layer:\n\nThe heading and the paragraph are ordinary Markdown. The tags are components from ChatGPT's catalog. The two `{@body …}` lines are JavaScript: the first declares a piece of state, `seats`, and the second derives `price` from it. The slider is bound to `seats`, so moving it updates the price. The full vocabulary is:\n\n- Markdown\n- component tags\n- `{@body …}` statements\n- `{expression}` interpolation\n- `{#if}` /`{:else if}` /`{:else}` conditionals\n- `{#each list as item, i}` loops\n- event handlers\n- `GenUI` actions such as`issueNewTurn` and`copy`\n\nA dedicated format is needed because the model writes the interface token by token.\n\n- **It has to be easy to write reliably,** so it is built from notation the model already knows well.\n- **It has to stay usable while half-written.** Statements sit on their own lines, and any open element can be closed automatically. That lets the server cut a partial response at its last complete construct and still compile it.\n\n## [Server-side compilation](#server-side-compilation)\n\nThe client never executes the model's output as written. The backend server compiles it into a JavaScript program and a JSON document, which are stored with the message (as `model_dil_v2`). The response compiles to this (formatted for readability):\n\nThe Markdown is compiled into the same tree as the components. The heading becomes a `title`, the paragraph a `text` with a `bold` inside it, and their words move into the constants table.\n\nCompilation does work that every client would otherwise have to repeat:\n\n- **Plain function calls.** Markup becomes calls to`__dil.jsx` , so a JavaScript runtime can evaluate the program without a parser for DIL.\n- **Error isolation.** Expressions are wrapped in`__dilSafe` , so an expression that throws removes one element instead of aborting the whole render.\n- **Text in a separate table.** Static text moves into the constants table, so as a response streams, growing text changes the data rather than the program.\n- **Stable state keys.** Each piece of state receives a key (`{ key: \"seats\" }` ), so its value survives every recompilation.\n- **Repair and validation.** Incomplete statements and tags are dropped, unclosed elements are closed, and properties that fail validation against the catalog are removed and recorded as diagnostics.\n\nThe JSON document holds the text constants and any data the server resolves for the response, such as image search results (see [Data](#data)).\n\n## [Client runtime](#client-runtime)\n\nThe client receives the compiled program and the JSON document. Its work is split between a runtime, which executes the program, and a renderer, which draws the result.\n\nThe program is model-written code, so it does not run in the ChatGPT page. ChatGPT loads a hidden iframe (`runner.html`), sandboxed with `allow-scripts` and a content security policy of `default-src 'none'`, which starts a Web Worker.\n\n- **Lockdown.** Before evaluating a program, the worker removes network access, timers, messaging, and dynamic code evaluation from its global scope, and freezes the remaining globals.\n- **Evaluation.** It then evaluates the program with`new Function` . The runtime objects (`DIL` ,`__dil` ,`GenUI` ) and the catalog's composite components are passed in as parameters.\n- **Watchdog.** A program that does not respond within a timeout is quarantined, and the worker is restarted.\n\nThe runtime is a small reconciler in the style of React. It renders the component and keeps hook state in keyed slots. It then compares the resulting tree with the previous one and encodes the differences as a list of operations. It does not draw anything.\n\nThe following illustrative example shows the operations from a first render, with one line per node. Entries that list each element's property names are omitted:\n\nFunctions never leave the worker; the slider's handler is sent only as an identifier (`fn#1`). On the wire, the operations are encoded as a binary sequence of integers, with strings held in a separate table.\n\n## [Rendering](#rendering)\n\nThe ChatGPT page applies the operations to its own component tree. Each `CREATE` instantiates a native component from ChatGPT's design system, and the page animates changes as they arrive. The page accepts operations only for known component types, so model output cannot introduce arbitrary markup or styles. The exceptions are the raw CSS values some properties accept (see [Design system and catalog](#design-system-and-catalog)) and `AppBlock` apps, which run in an iframe (see [AppBlock escape hatch](#appblock-escape-hatch)).\n\nInteraction runs in the opposite direction. When the user drags the slider to 9, the page sends the handler's identifier and arguments to the worker. The worker calls `setSeats(9)`, re-renders, and returns update operations.\n\nNo model call is involved.\n\nThe operation protocol does not depend on the platform. The worker's sandbox also permits the globals of Hermes, the JavaScript engine used by React Native. This suggests that ChatGPT's mobile apps run the same runtime and apply the operations with their own native renderers; we have not verified this directly.\n\n## [Design system and catalog](#design-system-and-catalog)\n\nThe catalog defines what the model can request. It is needed because the model does not build an interface from raw layout and styling rules. It chooses from components ChatGPT already knows how to draw, and styles them with design tokens such as `padding={3}`. Raw CSS values, such as pixel widths and hex colours, are accepted for some properties, but design tokens are preferred. As a result:\n\n- **Generated interfaces look like the rest of ChatGPT** on every platform.\n- **The compiler has a schema to check output against.** A property that does not exist on a component, or a literal of the wrong type, is removed during compilation and recorded as a diagnostic.\n\nIn one response we captured, the compiler removed two properties: `fill` on an `icon` (an `unknown_prop` diagnostic) and `gap=\"1\"` on a `box` (an `invalid_literal` diagnostic).\n\nThe catalog has three parts:\n\n- **Native components.** Around 70 components are defined in the component registry in ChatGPT's client code; 39 of them appear in the responses we captured.\n- **Design tokens** for spacing, radius, colour, and size.\n- **Composite components** written by OpenAI in DIL and sent to the sandbox prebuilt, such as the image and product components. In our captures, the model used these components but never defined its own.\n\nEvery part of the response maps to a catalog entry:\n\n| In the response | Catalog entry | Resolved value | \n|---|---|---|\n| `## Team plan estimate` | `title` | `lg` size token | \n| `**monthly price**` | `bold` inside`text` | Inline emphasis | \n| `box border padding={3} gap={2}` | Layout container | Border, 12 px padding, 8 px gap (4 px spacing scale) | \n| `slider min max value onChange` | Input | ChatGPT's slider | \n| `title size=\"xl\"` | Heading text | `xl` size token | \n\n## [Streaming](#streaming)\n\nStreaming text is simple: each new token is appended to what is already on screen. Streaming an interface is harder, for three reasons:\n\n- **The output is usually not runnable yet.** At most moments it is an incomplete program, with a tag or expression still open, and it cannot be executed as written.\n- **The interface has to keep working while it grows.** Components the user has already touched must keep their state.\n- **Some content arrives separately.** Data such as images comes from the server, not from the text.\n\nA plain stream of appended tokens cannot express this. ChatGPT instead streams patches to a structured message that holds the raw text, the compiled program, and its data side by side.\n\nThe response reaches the browser over a server-sent event stream (`POST /backend-api/f/conversation`). Each event is a JSON-Patch-style update to the message being built. A single event usually updates the raw DIL text and its compiled form together. This is one update from a captured response, shortened:\n\nThe server does not compile incrementally. Every few hundred milliseconds, most likely with each new chunk of model output, it recompiles everything the model has written so far and sends the result. Compilation starts with the first token, before any tag has appeared.\n\n### [Compiling a half-written response](#compiling-a-half-written-response)\n\nAt any moment, the model may be in the middle of a tag or an expression. The compiler cuts the source at its last complete construct:\n\n- an unfinished `{@body}` statement or tag is dropped;\n- an open element that already has content is closed automatically, and one without content is dropped;\n- partial text is kept as it is.\n\nThe compiler records each cut in `recoveryDiagnostics` (`unterminated_tag`, `unclosed_block`, `unterminated_braced_value`). The entry disappears once the source is complete again. In the interface, text streams word by word, while each component appears only once its tag is complete.\n\n### [Three kinds of update](#three-kinds-of-update)\n\nWe captured one response of 6,341 characters. Its text streamed over 18.7 seconds, in 84 updates. Of these, 76 carried new text, and they fell into three groups. The rest created the message (1), delivered image results (5), and marked completion (2).\n\n| Update | Count | What is sent | \n|---|---|---|\n| Structure changed | 52 | Appended text and the entire compiled program | \n| Only text grew | 14 | Appended text and an updated constant; the program is unchanged | \n| Inside an unfinished tag | 10 | Appended text only; the interface does not change | \n\nThe compiled program is a single nested expression whose closing brackets change with every recompile, so it cannot be appended to and is replaced in full. Re-sent code made up 83% of the roughly 275 KB of patches for this response. Updates arrived at a median interval of 230 ms, and full recompiles every 290 ms. In two other captures, updates arrived about every 410 ms.\n\n### [On the client](#on-the-client)\n\nThe page passes each new program to the sandboxed worker. The worker evaluates it, re-renders with the existing state, and sends update operations to the page. State keeps its values across recompiles because of the keys added during [compilation](#server-side-compilation). If a new program fails to evaluate or render, the worker keeps the last one that worked.\n\nThe page then animates each change:\n\n- text fades in over 0.7 s;\n- new rows and grid items slide in over 0.42 s;\n- charts draw over 1.8 s;\n- container heights transition instead of jumping.\n\n## [Data](#data)\n\nTool results contain values a user may act on: a retailer's price, a place's address, or an image from a page. Passing those values through the model risks copying errors or invented details. So ChatGPT supplies some of this data separately from the model's output: the model writes a request or a reference, and the server fills in the values alongside the compiled program, in `appData`. We saw two mechanisms for this, described below.\n\nNot all tool data takes this route. Weather figures from a web search, for example, were written into the response by the model itself.\n\n### [Server-defined components](#server-defined-components)\n\nSome components are resolved by the server. To show an image, the model describes it instead of linking to it:\n\nOnce the tag is complete, the server runs an image search, checks the resulting URLs, and patches the result into the response, typically a second or two later:\n\nIn the image-search responses we captured, image URLs were supplied by the server rather than written by the model. Some other components are resolved the same way:\n\n- `AsyncImageGroup` , for image carousels;\n- `Entity` , for product and place chips;\n- `Cite` , for source links.\n\n### [Binding tool results](#binding-tool-results)\n\nResults from the model's tool calls, such as a web search, are given IDs. The model refers to their fields by ID instead of retyping the values:\n\nThe server supplies the referenced fields with the response:\n\nThe price on screen comes from the search result itself, not from the model's copy of it.\n\n## [Actions](#actions)\n\nMost interactions never leave the response. Dragging a slider or ticking a checkbox changes state inside the worker, and the page receives only the resulting operations (see [Rendering](#rendering)). Actions are the interactions that reach outside the response. The program reaches them through a small `GenUI` object provided by the host, with functions like:\n\n- `issueNewTurn(text)`\n- `copy(text)`\n- `openUrl(url)`\n- `openEntityDetail(ref)`\n\n### [Continuing the conversation](#continuing-the-conversation)\n\nThe only way an interface communicates with the model is `issueNewTurn`. It sends a new user message, and the program builds that message's text from its state:\n\nAfter the user picked a style and a colour, the button produced this message:\n\nThe message is stored exactly like a typed one, with nothing marking it as coming from the interface. Forms work the same way: `<form onSubmit={…}>` together with `<button submit>` calls `issueNewTurn` with the form's values.\n\n## [AppBlock escape hatch](#appblock-escape-hatch)\n\nSome requests call for things the native components are not designed for, such as a drum machine that synthesizes sound with Web Audio. For these, the model can write an `AppBlock`: a self-contained web app in HTML, CSS, and JavaScript, embedded in the response. This is the start of one, shortened:\n\nAn `AppBlock` is handled differently from the rest of the response. The compiler does not translate it. The compiled program contains only a placeholder (`__chatgptClientDefinedWidget`), and the client reads the app's source from the raw response text.\n\nChatGPT renders the app in a visible iframe on a separate domain. It uses the same host (\"Skybridge\") that it uses for apps, and writes the model's HTML into an inner frame:\n\nAn app in an iframe sits outside ChatGPT's design system. Left alone, it would look like a foreign web page in the middle of the conversation. The host prevents this in three ways:\n\n- **Shared base styles.** It injects Tailwind's base styles, so the app starts from the same typography and spacing defaults as the page around it.\n- **Shared colours.** It provides theme variables (`--viz-text` ,`--viz-panel` ,`--viz-accent` and others) that the model's CSS refers to. The app uses these variables instead of hard-coded colours, so it follows ChatGPT's light and dark themes.\n- **Sizing.** It sizes the iframe to the app's content, so the app reads as part of the response rather than a scrolling box inside it.\n\nUnlike a native response, the app draws its own interface. Its state lives in its own JavaScript variables, and in our captures it made no calls back to ChatGPT.\n\n## [Markdown fallback](#markdown-fallback)\n\nA response is stored once, but it can be opened from clients that cannot run it, most likely older app versions. For these, the server generates a plain-Markdown version, `fallbackMarkdown`, alongside every compiled program. It keeps the prose. Images, products, citations, and maps are written in the inline markup that ChatGPT clients already render. Interactive parts are reduced to static text.\n\n## [Rough edges](#rough-edges)\n\n- **State resets on reload.** Values the user changes are lost when the page is reloaded, even though the client sends them to a view-state endpoint.\n- **The model does not see interface state.** It receives only the text of messages sent with`issueNewTurn` . When asked, it could not report values the user had changed.\n- **Follow-ups regenerate the interface.** Each follow-up produces a new response with a new program; the existing interface is never modified.\n- **Most of the stream is resent code.** In one response, resent programs made up 83% of 275 KB of patches for 6.3 KB of text.\n- **Interactions resend unchanged properties.** A single click in a pricing widget produced 158 operations, only 4 of which changed anything.\n- **Logic errors are silent.** An expression that throws an error renders nothing, and incorrect logic is not detected. In one dashboard, switching a chart's metric changed a revenue figure from $6,930 to $256,410.\n- **No apparent feedback on invalid properties.** The compiler removes invalid properties and records diagnostics, but the model repeated the same mistakes in later responses. This suggests that those diagnostics may not reach the model.\n- **The fallback omits computed content.** Values derived from state and content inside`{#each}` do not appear in the Markdown version.\n- **Image searches can miss.** The model writes a query but never sees the result. A query for a close-up of a handlebar shifter returned an image from an electric bicycle listing.\n- **Layouts are fixed.** Responses use fixed column counts and pixel widths, and none of the captured responses used the available breakpoint hooks.\n- **`AppBlock` apps are isolated.** Their state lives in their own JavaScript variables and is lost on reload, and they made no calls back to ChatGPT.\n\n## [Methodology](#methodology)\n\nAll observations come from our own ChatGPT accounts, from the traffic the ChatGPT web app generates, and from the JavaScript that chatgpt.com serves publicly. They were made in October 2026 with GPT-6 and GPT-6 Thinking.\n\n- **Conversation exports.** We exported conversations as the JSON that the web app loads for a conversation. For each response with an interface, the export holds the model's raw DIL, the compiled program, the constants and data, the Markdown fallback, and any compiler diagnostics.\n- **Stream captures.** We recorded the server-sent event stream for several responses by wrapping the page's stream reader in the browser. We then replayed the patches to rebuild the message after every update.\n- **Browser inspection.** We used the browser's developer tools to examine the page structure, including the iframes used for`AppBlock` and the scripts the worker evaluates.\n- **Client code.** We read the sandbox runner and worker (protocol version 14), and the chatgpt.com bundles that hold the component registry, the renderer, and the widget host. We then ran the unmodified runner in a local test page to record the exact operations it emits.\n- **Compiler reimplementation.** We reimplemented the server compiler from pairs of DIL source and compiled output. Its output matches OpenAI's byte for byte, both on the intermediate versions of a streamed response and on finished responses.\n- **Probes.** We wrote prompts designed to exercise specific features: forms, lists that can be edited, charts, maps, timers, and buttons that send messages back.\n\nMobile implementations were not examined.", "url": "https://wpnews.pro/news/we-reverse-engineered-chatgpt-intelligent-ui", "canonical_source": "https://www.openui.com/blog/how-chatgpt-intelligent-ui-works", "published_at": "2026-10-08 17:30:44+00:00", "updated_at": "2026-10-08 17:48:03.544595+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-products", "developer-tools"], "entities": ["OpenAI", "ChatGPT", "GPT-6", "GPT-6 Thinking", "DIL", "model_dil_v2", "__dilSafe", "__dil.jsx"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-reverse-engineered-chatgpt-intelligent-ui", "markdown": "https://wpnews.pro/news/we-reverse-engineered-chatgpt-intelligent-ui.md", "text": "https://wpnews.pro/news/we-reverse-engineered-chatgpt-intelligent-ui.txt", "jsonld": "https://wpnews.pro/news/we-reverse-engineered-chatgpt-intelligent-ui.jsonld"}}