{"slug": "how-we-check-every-ai-generated-room-render-before-a-user-sees-it", "title": "How we check every AI-generated room render before a user sees it", "summary": "A developer at AI Flip Room built a three-gate validation pipeline that checks every AI-generated room render before it reaches users, combining a structured-output vision judge, 16 indoor and 9 outdoor preservation checks, and a two-channel camera-shift measurement. The mechanical viewpoint detector caught 5 of 5 real camera shifts with 0 false flags on 24 hand-labelled pairs, while rewrites of the wall-art check reduced false failures from 10 of 16 good frames to 0 of 16 while still catching all 5 real portraits.", "body_md": "Image models are very good at restyling a room and quietly bad at leaving the room alone. Ask for a Japandi living room and you get beautiful furniture, and sometimes also:\n\nOur users put these images in property listings, where a staged photo may change the furniture but not the windows, walls or doors ([why that line exists](https://aifliproom.com/blog/virtual-staging-disclosure-rules)). So every render is checked automatically before anyone sees it. Here is how, and what building it taught us.\n\nFor an indoor room, each render goes through three gates in parallel:\n\nIf all three pass, the render is served. If not, we retry. If nothing passes, the user gets the best attempt, clearly marked as a draft.\n\nThe judge gets the original and the render, both downscaled, and must return a fixed JSON object. We use structured output with a schema, so it cannot reply in prose. An earlier version asked for \"only JSON\" and scraped it out with a regex; when the output got long or cut off, the parse failed and the gate silently did nothing.\n\nIndoors there are 16 checks, outdoors 9. Indoors they cover:\n\nOutdoors: buildings, fence line and gates, terrain (no new pool), mature trees, background, and a usable driveway.\n\nA simplified, illustrative verdict:\n\n```\n{\n  \"doors_preserved\": true,\n  \"openings_clear\": false,\n  \"windows_ok\": true,\n  \"window_unblocked\": true,\n  \"walls_uniform\": true,\n  \"frame_edges_solid\": true,\n  \"no_faces_in_art\": true,\n  \"fail_reasons\": [\"sofa end stands on the archway floor\"]\n}\n```\n\nA render passes only if every counted boolean is true; the reasons feed logs and the retry prompt. Most of the work went into the wording:\n\n**Say what must pass.** Restyling is the product, so the instructions list what is expected (new floors, paint, wallpaper, fixtures) and say \"unsure means pass\". The exception is a short list where unsure means fail, led by a window that may not have been in the original. An invented window is the worst defect a listing photo can have.\n\n**One question per field.** While the frame-edge rule was a sub-clause of a broader \"no phantom architecture\" check, new doorways at the frame edge got through. Its own boolean made the model actually look.\n\n**Naming things primes the judge.** An early niche check listed which shelves were allowed, and the judge started flagging every shelf. The current wording is purely geometric: is this flat wall still one plane?\n\n**Allow \"nothing to compare\".** The \"same view through the window\" check kept firing on originals with blinds drawn or glass blown out to white. That case is now an explicit pass.\n\n**Wording is measurable.** The old wall-art wording wrongly failed 10 of 16 good frames (stylised line drawings counted as faces). The rewrite failed 0 of 16 and still caught all 5 real portraits.\n\nEarly on, the vision model was close to blind to viewpoint changes: renders shot from a visibly different spot passed every neural check. Later, a stronger judge had the opposite problem: on 43 renders where the camera had not moved, it said it had 13 times. So indoors the judge's camera answers no longer decide pass or fail. A measurement does, with two channels:\n\nEach channel is fooled by something different. Paint fools the ceiling line (dark beams, art on a mantel). Glass fools depth (windows and mirrors are estimated differently each time). Paint is flat in depth, and glass does not move the ceiling line, so we flag a shift only when **both** agree: the minimum of the two normalised scores.\n\nOn 24 hand-labelled pairs from two rooms, this caught 5 of 5 real shifts with 0 false flags. The margin was thin, so a grey zone around the threshold is logged, not failed. Outdoors there is no ceiling to trace, so the mechanical check is off and the judge's camera question stays in charge.\n\nAt the judge's resolution a front door is a thin strip. So doors, windows and stairs are detected at upload, and after each render a model is asked about each crop at full resolution: is it still there? A mechanical channel also measures whether a window drifted or widened. On 87 labelled pairs, the crop judge caught 7 of 10 broken renders with no false alarms on the 75 clean ones; the mechanical channel caught the other 3.\n\nRetries depend on the kind of defect.\n\nThe geometry retry gets a fixed, count-based instruction (\"same set of openings, not one more, not one less\"), never the judge's description of the defect. When a retry prompt named an arched opening at the left edge, the retry model painted exactly that.\n\nIf nothing passes, the user still gets the best attempt, under a \"rough draft\" banner. On paid plans the first four rejected drafts in a rolling week don't use a render; the free plan gets one. The cap is deliberate: a draft is a real, downloadable image, so unlimited forgiveness would be a free image farm.\n\nThe checks fail open. If the judge itself errors or times out, the user is not blocked: the image is served flagged, and an unverified render is not charged as a draft.\n\nAt first, \"best attempt\" meant the one with the fewest failed checks. The attempt that breaks the least is often the one where the model barely did anything. In one real case the first attempt was a convincing industrial restyle that failed on geometry, and the user was served the second: their old room, old wallpaper and all, plus a pair of lamps. The ranking was paying the model for doing nothing.\n\nThe fix is a change meter: shrink both images to a tiny thumbnail and count how much of the frame visibly changed. No model call, and we calibrated it on 110 real renders and 10 synthetic near-copies (recompressed, blurred, brightened, shifted, a couple of lamp-sized patches). The ranking now:\n\n```\n// simplified\nfunction draftRank(a: Attempt): number {\n  if (a.checkerErrored) return LAST;\n  let rank = a.failedChecks.length;\n  if (a.hasPeopleOrRealFacesInArt) rank += HARD_BAN;\n  if (a.isNearCopy && a.intensity !== \"light\") rank += NEAR_COPY;\n  return rank; // lowest is served\n}\n```\n\nA light restage is exempt, because there the user asked us to keep the room.\n\nThe judge runs on every attempt, so a cheaper model is tempting. We keep a calibration set of 31 hand-labelled frames and run every candidate on it twice. A cheaper model at under half the price per call missed 6 defects per run against our current judge's 3. On 46 live renders with 8 defects marked by eye, it caught 0 and 1 across two runs; the current judge caught 6. A newer model wasn't clearly better either: fewer misses (2 vs 3), twice the false alarms (4 vs 2). We kept the current judge.\n\nWe build this at [AI Flip Room](https://aifliproom.com).", "url": "https://wpnews.pro/news/how-we-check-every-ai-generated-room-render-before-a-user-sees-it", "canonical_source": "https://dev.to/aifliproom/how-we-check-every-ai-generated-room-render-before-a-user-sees-it-3ek8", "published_at": "2026-10-03 02:05:35+00:00", "updated_at": "2026-10-03 02:07:47.130108+00:00", "lang": "en", "topics": ["computer-vision", "generative-ai", "ai-tools", "ai-products", "structured-data"], "entities": ["AI Flip Room"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-we-check-every-ai-generated-room-render-before-a-user-sees-it", "markdown": "https://wpnews.pro/news/how-we-check-every-ai-generated-room-render-before-a-user-sees-it.md", "text": "https://wpnews.pro/news/how-we-check-every-ai-generated-room-render-before-a-user-sees-it.txt", "jsonld": "https://wpnews.pro/news/how-we-check-every-ai-generated-room-render-before-a-user-sees-it.jsonld"}}