My coding agent could already render a single Android XML layout through Robolectric and get back a PNG plus a View Tree with pixel bounds. Then I pointed it at a real screen from a work app, and one layout stopped being enough.
A real screen is an Activity shell, a Fragment container, a RecyclerView with data, and states like a overlay. Launching the production Fragment would drag in DI, navigation, networking and the rest of the app's runtime. I wanted something narrower: a screen real enough for visual checks, without running the app.
This is how android-ui-renderer-mcp got composed-screen rendering.
The work screen can't be shown, so every image here is a render of an open demo app in the same repository (sample/). It has the same structure — Activity with a toolbar, Fragment, RecyclerView, master-detail, overlay — and each image can be reproduced from the request stored next to it.
Activity XML — real
Fragment XML — real
item layout — real
drawable — real
Fragment runtime — not run
DI — not run
navigation — not run
network — not run
data — set explicitly
That is the main design decision: the renderer uses the app's real UI resources, but takes runtime state from a deterministic request. Otherwise it becomes one more way to run the app, only harder than an emulator.
A new render_target tool supports one target kind, activity_fragment. It inflates the Activity XML, finds the container, inflates the Fragment XML separately and inserts it. The production Fragment class never starts.
{
"target": {
"kind": "activity_fragment",
"activityLayout": "activity_main",
"containerId": "fragmentContainer",
"fragmentLayout": "fragment_library"
},
"widthPx": 1920,
"heightPx": 1200,
"densityDpi": 240,
"orientation": "landscape"
}
The toolbar title and menu icons come from XML (app:title, app:menu); no Activity code ran. The details on the right are an <include>, filled from the request.
I'm not emulating the Fragment lifecycle. I need its visual result from real resources, because the agent changes XML and has to see what ends up on screen.
An empty RecyclerView says little about the real UI, and running the production adapter would pull real data back in. So list data became part of the request: the real itemLayout, the orientation, the rows, and a fixture per row. The renderer creates a temporary adapter and inflates the real item XML for each row.
One row from the request above, the selected book:
{
"@id/title": { "text": "Designing Data-Intensive Applications" },
"@id/cover": { "image": { "type": "drawable_resource", "value": "@drawable/cover_green" } },
"@id/statusBadge": { "text": "On loan", "backgroundDrawable": "@drawable/bg_badge_on_loan", "textColor": "#8A4B08" },
"@id/root": { "selected": true, "backgroundDrawable": "@drawable/bg_book_row_selected" }
}
The UI structure is real; the data is test data and fully controlled. The renderer doesn't guess values from tools:*. If the agent wants three rows, it describes three rows.
A overlay above Activity + Fragment is just another XML layout with its own fixture. The catch is the indeterminate ProgressBar: it animates, so two screenshots in a row can differ. For automated visual checks that's needless nondeterminism, so the renderer freezes it into a static ring before drawing.
This doesn't prove the animation works. It proves the indicator exists, sits in the right place and has the expected geometry, which is the level of realism this task needs.
For the work screen I also pinned down the real tablet profile: 1280×800 screen, 1280×728 app area (72 px at the bottom taken by system navigation), densityDpi 240, fontScale 1.0, ru-RU, day mode, landscape. For a screenshot that's a detail; for comparing bounds it isn't. The agent has to measure in the same area the real screen has.
At some point I opened a saved render and realized its PNG and View Tree were there, but the call that produced them wasn't: which tool, which arguments, which project root, module, variant and Gradle test task. A reproducible renderer with a non-reproducible run history.
Every sidecar run now also saves request.json (the normalized request) and replay.json (the recipe). For the overlay screen above, shortened:
{
"format": "android-ui-renderer-mcp/replay/v1",
"tool": "render_target",
"arguments": {
"target": { "kind": "activity_fragment", "activityLayout": "activity_main", "containerId": "fragmentContainer", "fragmentLayout": "fragment_library" },
"widthPx": 1920, "heightPx": 1200, "densityDpi": 240, "orientation": "landscape",
"overlays": [{ "layout": "view__overlay", "fixture": { "@id/Root": { "visibility": "visible" } } }]
},
"project": { "root": "sample", "module": ":app", "variant": "debug", "testTask": ":app:testDebugUnitTest" }
}
A render session is now a small experiment you can repeat on the same project revision and renderer version.
I built the demo later, only to have images I could publish. It found real issues.
The renderer silently required JUnit. The first demo render failed:
error: package org.junit does not exist
The probe is a JUnit 4 test, but the init script added only Robolectric. The work project had JUnit in its test dependencies, like any project from the default Android Studio template, so this never showed up. Now JUnit 4.13.2 is added for the render run only if the project doesn't declare it; a project on JUnit 4.12 stays on 4.12, checked against the test classpath.
A night-mode contrast bug in the demo itself. An outlined button was almost invisible because the theme's primary color matched the dark toolbar. The nightMode: true render showed it; reading the XML didn't.
Byte-identical re-renders. Re-running all seven demo scenarios on the same machine produced PNGs identical to the saved ones. With replay.json, a replay gives the same image, not a similar one.
The View Tree sees what the PNG hides. A long title at fontScale: 1.3:
"textLayout": { "textSizePx": 40.0, "maxLines": 1, "lineCount": 1, "ellipsisCount": 81, "truncated": true }
The bounds didn't change and the image looks fine, but 81 characters became an ellipsis. Also, 16sp at fontScale: 1.3 is 40 px, not 41.6: since Android 14, large text scales non-linearly.
I have a detailed worklog for this iteration.
Me: picked the real screen as a testing ground and made the work app read-only; asked for RecyclerView support; supplied a mockup as the source of test data and the real tablet profile; stopped the agent before the overlay/dialog changes and asked for analysis first; chose a separate render_target over over render_layout; noticed that renders couldn't be replayed; checked the composed render with the overlay myself before allowing the commit. For the demo, I defined the screens and asked for JUnit to be added only when missing.
The coding agent: studied the real screen's XML and the renderer's limits, proposed the deterministic RecyclerView adapter, implemented list and drawable-state support, the composed target, overlays and the replay recipe, ran checks on an isolated copy of the work project without changing it, and wrote the demo app, its render requests and the JUnit fix.
I didn't type most of the code. I decided which part of reality the tool should model and what it must not do.
activity_fragment: one container, one fragment layout. Master-detail in the demo is a single fragment layout, not two fragments.
Each step of this project came down to one question: which part of the real app must stay real, and which can be replaced with a deterministic model? My current answer: resources and geometry stay real; data and runtime state are explicit and controlled; and every render leaves enough behind to be repeated.
The renderer and the demo are open source: hram/android-ui-renderer-mcp. The full case study is on my site: https://hram.github.io/en/articles/android-ui-renderer-composed-screens/