cd /news/artificial-intelligence/dags-disentangled-appearance-and-geo… · home › topics › artificial-intelligence › article
[ARTICLE · art-145156] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering

Researchers introduced DAGS, a lightweight, attention-free conditioning scheme that steers a frozen image diffusion transformer (DiT) to produce high-fidelity, independently controllable renders, according to the arXiv paper 2610.02567v1. On a matched 1-spp + G-buffer input, DAGS reconstructs +8.6 dB and +10.1 dB PSNR over the real-time denoiser Intel OIDN and the diffusion renderer RGB<->RGB respectively, while being 2.5-8x more temporally stable perceptually as measured by temporal-LPIPS flicker. DAGS is not real-time, trading compute for controllability and quality, and uses two small convolutional encoders plus a recurrent lighting stabilizer and a training-free temporal guidance term to turn a per-frame image model into a streaming renderer.

by read1 min views1 publishedOct 5, 2026

arXiv:2610.02567v1 Announce Type: new Abstract: Diffusion transformers (DiTs) generate high-fidelity images from text and image conditions, but their outputs carry large variance and their faithfulness to a desired target depends heavily on how the condition is supplied. We present DAGS, a lightweight, attention-free, disentangled appearance and geometry conditioning scheme that steers a frozen image DiT to produce high-fidelity, highly faithful, and independently controllable renders. Two small convolutional encoders compute conditioning features once per frame and inject them as a learned, per-layer, element-wise residual into the image tokens, avoiding the quadratic cost of stacking conditions through attention. Because control and temporal handling live outside the frozen backbone, we retain its vast pretrained prior and eliminate backbone-overfitting risk. We further add a small recurrent lighting stabilizer and a training-free temporal guidance term that, coupled with our conditioning, elevate a per-frame image model into a streaming renderer. DAGS produces controllable, high-quality renders at a fraction of the compute of path tracing; it is not real-time, trading compute for controllability and quality. On a matched 1-spp + G-buffer input, per-frame DAGS reconstructs +8.6 dB / +10.1 dB PSNR over the real-time denoiser Intel OIDN and the diffusion renderer RGB<->X while being 2.5-8x more temporally stable perceptually (temporal-LPIPS flicker).

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @dags 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dags-disentangled-ap…] indexed:0 read:1min 2026-10-05 · —