How Diffusion Controller unifies and simplifies AI image generation Google Research software engineers Chih-wei Hsu and Moonkyung Ryu introduced Diffusion Controller, a lightweight "steering damper" network that attaches to text-to-image models, including access-restricted closed-source ones, to improve prompt alignment without breaking baseline stability. The fully unlocked version, with unrestricted access to alter internal model weights, achieved a 90% win rate over the baseline model, and the lightweight add-on network outperformed the industry standard for matching human preferences. The framework reframes the denoising process as a continuous control problem, unifying inference-time guidance techniques such as classifier-free diffusion guidance with fine-tuning approaches like LoRA, reward-weighted regressions and policy gradients. September 29, 2026 Chih-wei Hsu and Moonkyung Ryu , Software Engineers, Google Research We introduce Diffusion Controller, a lightweight "steering damper" network that precisely steers image generation to achieve significantly better prompt alignment. It seamlessly attaches to even access-restricted, closed-source models, boosting image quality without breaking baseline stability. The rapid advancement of text-to-image AI models, such as Nano Banana https://gemini.google/overview/image-generation/ , Stable Diffusion https://openart.ai/image/create/stable-diffusion-3-sd3?utm source=google&utm medium=cpc&utm campaign=search-google-acq-web na search ai models&utm content=GPT image&utm term=gpt%20image2&utm campaign=NA Search AI Models AI Max LTV&utm source=adwords&utm medium=ppc&utm id=23701882222&hsa acc=9574504800&hsa cam=23701882222&hsa grp=201400335528&hsa ad=806159705184&hsa src=g&hsa tgt=kwd-2481063683915&hsa kw=gpt%20image2&hsa mt=e&hsa net=adwords&hsa ver=3&gad source=1&gad campaignid=23701882222&gbraid=0AAAAAp6wzGRqGO35lRpoSrG3vkuAFWuo2&gclid=CjwKCAjwqJXUBhBNEiwA8BgG7sNjtkqRlaQVQpNWJdtcU3xJ2Y2gCXkdnZCLjcDLMBvRsPfIrwl1FhoC7g4QAvD BwE and Flux https://elevenlabs.io/image/flux2-pro?utm source=google&utm medium=cpc&utm campaign=na nonbrandsearch image-video english&utm id=23646367622&utm term=flux%20ai%20generator&utm content=image video - ai models - flux&gad source=1&gad campaignid=23646367622&gbraid=0AAAAAp9ksTGNovuz5 8YDLVdPu5FBbulO&gclid=CjwKCAjwqJXUBhBNEiwA8BgG7rX-GgR6YWQYihPEvcS7rXRbRAqpV3q05Ceze8sSUkYbIsvL8B4HshoCKrQQAvD BwE , has fundamentally transformed creative design, allowing anyone to synthesize photorealistic, high-fidelity images from textual descriptions. However, steering these massive models to meet precise user intent, downstream goals, or strict visual constraints remains a delicate and unpredictable balancing act. For example, imagine prompting a model for "a lizard wearing sunglasses". The model might generate a realistic lizard that's not wearing sunglasses. Alternatively, forcing the model to include the sunglasses might distort the lizard's face, ruining the image quality. Existing methodologies that guide or fine-tune image generation are very disconnected. On the one hand, developers use inference-time techniques https://magazine.sebastianraschka.com/p/categories-of-inference-time-scaling e.g., classifier-free diffusion guidance https://arxiv.org/abs/2207.12598 to adjust the text prompt’s influence and guide the image generation process on the fly. On the other hand, they rely on heavy fine-tuning