September 29, 2026
Chih-wei Hsu and Moonkyung Ryu , Software Engineers, Google Research
We introduce Diffusion Controller, a lightweight "steering damper" network that precisely steers image generation to achieve significantly better prompt alignment. It seamlessly attaches to even access-restricted, closed-source models, boosting image quality without breaking baseline stability.
The rapid advancement of text-to-image AI models, such as Nano Banana, Stable Diffusion and Flux, has fundamentally transformed creative design, allowing anyone to synthesize photorealistic, high-fidelity images from textual descriptions. However, steering these massive models to meet precise user intent, downstream goals, or strict visual constraints remains a delicate and unpredictable balancing act. For example, imagine prompting a model for "a lizard wearing sunglasses". The model might generate a realistic lizard that's not wearing sunglasses. Alternatively, forcing the model to include the sunglasses might distort the lizard's face, ruining the image quality.
Existing methodologies that guide or fine-tune image generation are very disconnected. On the one hand, developers use inference-time techniques (e.g., classifier-free diffusion guidance) to adjust the text prompt’s influence and guide the image generation process on the fly. On the other hand, they rely on heavy fine-tuning using parameter-efficient adapters like LoRA, reward-weighted regressions, or policy gradients to alter a model's behavior.
Because these tools have historically been treated as distinct and unrelated fixes, the field has lacked a single, principled mathematical language to unify, analyze, and optimize how we control generative models. This fragmented approach often forces engineers to rely on guesswork when balancing user preference alignment against image quality.
To solve this balancing act, we present the Diffusion Controller framework. Instead of treating image generation as a rigid sequence of isolated steps, Diffusion Controller reframes the entire denoising process as a smooth, continuous control problem. Our results show that Diffusion Controller’s lightweight add-on network outperformed the industry standard for matching human preferences. Moreover, its fully unlocked version (i.e., the fine-tuned model with "white-box" or unrestricted access to alter internal model weights) achieved a 90% win rate over the baseline model.
The Diffusion Controller framework treats the image generation process (where an AI model starts with random noise and gradually refines it into a clear picture) as a smoothly controlled journey. Imagine the base pre-trained model as a massive, powerful motorcycle; rebuilding its core engine to change how it drives is inefficient and risky. Instead, the Diffusion Controller acts as a lightweight steering damper attached to it while the main model (the motorcycle) remains completely frozen and safely untouched.
Rather than guessing how to guide the generation at each step, the Diffusion Controller’s core mechanism (the steering damper) dynamically adjusts the generation trajectory as the image is created. It smoothly recalibrates the model’s standard, default behavior, giving more weight to directions that maximize a user-defined target (such as achieving an artistic style or contextual alignment).
By mathematically optimizing these shifts in the generation trajectory with feedback, the Diffusion Controller framework strikes a balance, successfully steering the model’s generation process toward new user preferences while fully preserving the base model’s crisp visual image quality and underlying stability. Returning to the lizard example above, this means the steering damper steers the model to ensure the sunglasses are included, but the penalty guardrail kicks in to prevent the system from distorting the lizard's natural scales and proportions to make it happen.
To turn this theoretical control problem into a practical tool, we bridged the gap between abstract and possibly intractable equations into two efficient fine-tuning methods based entirely on a final reward score:
The Diffusion Controller framework operates like a steering damper, which solves a massive real-world business problem. Usually, to change how a model behaves, you need "white-box" access to dig into its core engine and change its internal settings. However, the world’s best image-generation models are often corporate secrets (referred to as black boxes or gray boxes).
The steering damper network relies on the perfect image-steering instruction, a combination of the base model’s knowledge plus a small correction. It observes the image as it is being cleared of static, and injects precise, microscopic steering corrections. This allows engineers to perfectly control and customize even tightly locked, closed-source models without ever touching the underlying code.
We evaluated the Diffusion Controller framework’s capabilities using a Stable Diffusion v1.4 backbone across three fine-tuning regimes: supervised fine-tuning (SFT), reward-weighted loss (RWL), and PPO. Performance was measured using the standardized Human Preference Score (HPS-v2) to determine how well generated images aligned with user prompts and aesthetic choices.
We have implemented four network structures under the Diffusion Controller framework:
The results demonstrated that Diffusion Controller excels in both fully accessible "white-box" environments and highly restrictive "gray-box" environments, consistently outperforming its corresponding baselines and achieving a better quality-efficiency trade-off.
In the SFT and RWL tracks, the gray-box Diffusion Controller steering damper network outperformed LoRA — the state-of-the-art parameter-efficient, white-box approach — in HPPS-v2 win rates. This is a significant milestone, considering Diffusion Controller accomplished this while manipulating significantly fewer internal model layers than LoRA. Furthermore, in comprehensive human evaluation panels, Diffusion Controller recorded the best subjective quality and prompt-matching results across complex, multi-attribute test prompts.
A core feature of the framework is its flexibility at runtime. By adjusting a single inference-time guidance strength parameter, users can dynamically dial up or down the intensity of the control constraints. This allows for smooth, granular adjustment of prompt alignment on the fly without breaking baseline image stability or causing the visual distortions typical of older guidance methods.
By combining a mathematical control with an image generation model, Diffusion Controller bridges the gap between pure mathematics and modern creative tools. Instead of relying on a patchwork of guesswork and quick fixes to tune a model, it provides a single, mathematically sound system for steering image generation. Best of all, it offers a practical, lightweight "steering damper" blueprint that works perfectly even on access-restricted, closed-source models.
Looking ahead, this unified framework opens up exciting new paths for research. Because the control layer is completely separated from the model's core engine, future developers can use Diffusion Controller for much more than just matching text prompts. Immediate next steps include using the framework to build advanced personalization tools, developing robust safety mechanisms to help mitigate harmful content generation, and adapting the steering damper network to control complex, next-generation video models.
We would like to thank our co-authors and collaborators from Google Research, Google DeepMind, and academia for their contributions to this work.