We reverse engineered ChatGPT Intelligent UI A technical analysis published in October 2026 details how ChatGPT's Intelligent UI, introduced with GPT-6, renders interactive components inline in conversations by having the model write interfaces in a language OpenAI calls DIL, which combines Markdown, JSX-like tags, and JavaScript. The backend compiles each partial response into a JavaScript program and a JSON document stored as model_dil_v2, wrapping expressions in __dilSafe for error isolation and assigning stable state keys so values survive recompilation, while the client executes the program in a sandbox and applies the resulting operations to ChatGPT's native components. The format is designed to be writable token by token and usable while half-written, letting the server cut a partial response at its last complete construct and still compile it. What is Intelligent UI? what-is-intelligent-ui Intelligent UI is the capability, introduced with GPT-6, that allows ChatGPT to respond with interactive interfaces rendered inline in the conversation OpenAI announcement https://openai.com/index/gpt-6-for-everyone/ . A single response can combine prose with components such as sliders, forms, tables, charts, maps, and product cards. These components respond to input without a further model call, and they are drawn with ChatGPT's own design system rather than embedded as a web page. The observations in this article were made on ChatGPT for the web in October 2026, using GPT-6 and GPT-6 Thinking. Building blocks building-blocks ChatGPT's implementation divides the work between the model, the backend server, and the client: - Inference format inference-format : the model writes the interface in DIL, which combines Markdown with JSX-like tags and JavaScript. - Server-side compilation server-side-compilation : the server converts each partial response into a JavaScript program and a JSON document of text and data. - Client runtime client-runtime : a sandboxed runtime executes the program and produces UI operations. - Rendering rendering : ChatGPT applies the operations to its own native components. - Design system and catalog design-system-and-catalog : the components, properties, and design tokens available to the model. Inference format inference-format This is what the model writes. In ChatGPT, it is a language OpenAI calls DIL: Markdown for prose, JSX-like tags for components, and JavaScript for state and logic. We will follow one small response through every layer: The heading and the paragraph are ordinary Markdown. The tags are components from ChatGPT's catalog. The two {@body …} lines are JavaScript: the first declares a piece of state, seats , and the second derives price from it. The slider is bound to seats , so moving it updates the price. The full vocabulary is: - Markdown - component tags - {@body …} statements - {expression} interpolation - { if} / {:else if} / {:else} conditionals - { each list as item, i} loops - event handlers - GenUI actions such as issueNewTurn and copy A dedicated format is needed because the model writes the interface token by token. - It has to be easy to write reliably, so it is built from notation the model already knows well. - It has to stay usable while half-written. Statements sit on their own lines, and any open element can be closed automatically. That lets the server cut a partial response at its last complete construct and still compile it. Server-side compilation server-side-compilation The client never executes the model's output as written. The backend server compiles it into a JavaScript program and a JSON document, which are stored with the message as model dil v2 . The response compiles to this formatted for readability : The Markdown is compiled into the same tree as the components. The heading becomes a title , the paragraph a text with a bold inside it, and their words move into the constants table. Compilation does work that every client would otherwise have to repeat: - Plain function calls. Markup becomes calls to dil.jsx , so a JavaScript runtime can evaluate the program without a parser for DIL. - Error isolation. Expressions are wrapped in dilSafe , so an expression that throws removes one element instead of aborting the whole render. - Text in a separate table. Static text moves into the constants table, so as a response streams, growing text changes the data rather than the program. - Stable state keys. Each piece of state receives a key { key: "seats" } , so its value survives every recompilation. - Repair and validation. Incomplete statements and tags are dropped, unclosed elements are closed, and properties that fail validation against the catalog are removed and recorded as diagnostics. The JSON document holds the text constants and any data the server resolves for the response, such as image search results see Data data . Client runtime client-runtime The client receives the compiled program and the JSON document. Its work is split between a runtime, which executes the program, and a renderer, which draws the result. The program is model-written code, so it does not run in the ChatGPT page. ChatGPT loads a hidden iframe runner.html , sandboxed with allow-scripts and a content security policy of default-src 'none' , which starts a Web Worker. - Lockdown. Before evaluating a program, the worker removes network access, timers, messaging, and dynamic code evaluation from its global scope, and freezes the remaining globals. - Evaluation. It then evaluates the program with new Function . The runtime objects DIL , dil , GenUI and the catalog's composite components are passed in as parameters. - Watchdog. A program that does not respond within a timeout is quarantined, and the worker is restarted. The runtime is a small reconciler in the style of React. It renders the component and keeps hook state in keyed slots. It then compares the resulting tree with the previous one and encodes the differences as a list of operations. It does not draw anything. The following illustrative example shows the operations from a first render, with one line per node. Entries that list each element's property names are omitted: Functions never leave the worker; the slider's handler is sent only as an identifier fn 1 . On the wire, the operations are encoded as a binary sequence of integers, with strings held in a separate table. Rendering rendering The ChatGPT page applies the operations to its own component tree. Each CREATE instantiates a native component from ChatGPT's design system, and the page animates changes as they arrive. The page accepts operations only for known component types, so model output cannot introduce arbitrary markup or styles. The exceptions are the raw CSS values some properties accept see Design system and catalog design-system-and-catalog and AppBlock apps, which run in an iframe see AppBlock escape hatch appblock-escape-hatch . Interaction runs in the opposite direction. When the user drags the slider to 9, the page sends the handler's identifier and arguments to the worker. The worker calls setSeats 9 , re-renders, and returns update operations. No model call is involved. The operation protocol does not depend on the platform. The worker's sandbox also permits the globals of Hermes, the JavaScript engine used by React Native. This suggests that ChatGPT's mobile apps run the same runtime and apply the operations with their own native renderers; we have not verified this directly. Design system and catalog design-system-and-catalog The catalog defines what the model can request. It is needed because the model does not build an interface from raw layout and styling rules. It chooses from components ChatGPT already knows how to draw, and styles them with design tokens such as padding={3} . Raw CSS values, such as pixel widths and hex colours, are accepted for some properties, but design tokens are preferred. As a result: - Generated interfaces look like the rest of ChatGPT on every platform. - The compiler has a schema to check output against. A property that does not exist on a component, or a literal of the wrong type, is removed during compilation and recorded as a diagnostic. In one response we captured, the compiler removed two properties: fill on an icon an unknown prop diagnostic and gap="1" on a box an invalid literal diagnostic . The catalog has three parts: - Native components. Around 70 components are defined in the component registry in ChatGPT's client code; 39 of them appear in the responses we captured. - Design tokens for spacing, radius, colour, and size. - Composite components written by OpenAI in DIL and sent to the sandbox prebuilt, such as the image and product components. In our captures, the model used these components but never defined its own. Every part of the response maps to a catalog entry: | In the response | Catalog entry | Resolved value | |---|---|---| | Team plan estimate | title | lg size token | | monthly price | bold inside text | Inline emphasis | | box border padding={3} gap={2} | Layout container | Border, 12 px padding, 8 px gap 4 px spacing scale | | slider min max value onChange | Input | ChatGPT's slider | | title size="xl" | Heading text | xl size token | Streaming streaming Streaming text is simple: each new token is appended to what is already on screen. Streaming an interface is harder, for three reasons: - The output is usually not runnable yet. At most moments it is an incomplete program, with a tag or expression still open, and it cannot be executed as written. - The interface has to keep working while it grows. Components the user has already touched must keep their state. - Some content arrives separately. Data such as images comes from the server, not from the text. A plain stream of appended tokens cannot express this. ChatGPT instead streams patches to a structured message that holds the raw text, the compiled program, and its data side by side. The response reaches the browser over a server-sent event stream POST /backend-api/f/conversation . Each event is a JSON-Patch-style update to the message being built. A single event usually updates the raw DIL text and its compiled form together. This is one update from a captured response, shortened: The server does not compile incrementally. Every few hundred milliseconds, most likely with each new chunk of model output, it recompiles everything the model has written so far and sends the result. Compilation starts with the first token, before any tag has appeared. Compiling a half-written response compiling-a-half-written-response At any moment, the model may be in the middle of a tag or an expression. The compiler cuts the source at its last complete construct: - an unfinished {@body} statement or tag is dropped; - an open element that already has content is closed automatically, and one without content is dropped; - partial text is kept as it is. The compiler records each cut in recoveryDiagnostics unterminated tag , unclosed block , unterminated braced value . The entry disappears once the source is complete again. In the interface, text streams word by word, while each component appears only once its tag is complete. Three kinds of update three-kinds-of-update We captured one response of 6,341 characters. Its text streamed over 18.7 seconds, in 84 updates. Of these, 76 carried new text, and they fell into three groups. The rest created the message 1 , delivered image results 5 , and marked completion 2 . | Update | Count | What is sent | |---|---|---| | Structure changed | 52 | Appended text and the entire compiled program | | Only text grew | 14 | Appended text and an updated constant; the program is unchanged | | Inside an unfinished tag | 10 | Appended text only; the interface does not change | The compiled program is a single nested expression whose closing brackets change with every recompile, so it cannot be appended to and is replaced in full. Re-sent code made up 83% of the roughly 275 KB of patches for this response. Updates arrived at a median interval of 230 ms, and full recompiles every 290 ms. In two other captures, updates arrived about every 410 ms. On the client on-the-client The page passes each new program to the sandboxed worker. The worker evaluates it, re-renders with the existing state, and sends update operations to the page. State keeps its values across recompiles because of the keys added during compilation server-side-compilation . If a new program fails to evaluate or render, the worker keeps the last one that worked. The page then animates each change: - text fades in over 0.7 s; - new rows and grid items slide in over 0.42 s; - charts draw over 1.8 s; - container heights transition instead of jumping. Data data Tool results contain values a user may act on: a retailer's price, a place's address, or an image from a page. Passing those values through the model risks copying errors or invented details. So ChatGPT supplies some of this data separately from the model's output: the model writes a request or a reference, and the server fills in the values alongside the compiled program, in appData . We saw two mechanisms for this, described below. Not all tool data takes this route. Weather figures from a web search, for example, were written into the response by the model itself. Server-defined components server-defined-components Some components are resolved by the server. To show an image, the model describes it instead of linking to it: Once the tag is complete, the server runs an image search, checks the resulting URLs, and patches the result into the response, typically a second or two later: In the image-search responses we captured, image URLs were supplied by the server rather than written by the model. Some other components are resolved the same way: - AsyncImageGroup , for image carousels; - Entity , for product and place chips; - Cite , for source links. Binding tool results binding-tool-results Results from the model's tool calls, such as a web search, are given IDs. The model refers to their fields by ID instead of retyping the values: The server supplies the referenced fields with the response: The price on screen comes from the search result itself, not from the model's copy of it. Actions actions Most interactions never leave the response. Dragging a slider or ticking a checkbox changes state inside the worker, and the page receives only the resulting operations see Rendering rendering . Actions are the interactions that reach outside the response. The program reaches them through a small GenUI object provided by the host, with functions like: - issueNewTurn text - copy text - openUrl url - openEntityDetail ref Continuing the conversation continuing-the-conversation The only way an interface communicates with the model is issueNewTurn . It sends a new user message, and the program builds that message's text from its state: After the user picked a style and a colour, the button produced this message: The message is stored exactly like a typed one, with nothing marking it as coming from the interface. Forms work the same way: