Setup Diffusers for Stable Diffusion XL Using a Nvidia Pascal GPU in Linux A developer documented a working Linux setup for running Stable Diffusion XL through Hugging Face Diffusers on an Nvidia GTX 1070 Pascal GPU, pinning PyTorch to the CUDA 12.4 wheel index and transformers to version 5.5.0. The guide states that 5.5.0 was the last transformers release compatible with newer Diffusers versions, and provides a Python script using StableDiffusionXLPipeline with DPMSolverMultistepScheduler, 26 sampling steps, a 6.5 guidance scale, and bfloat16 precision to generate two images per prompt. You need to first ensure that all the necessary drivers CUDA most important for the Nvidia GPU are installed. This post will NOT cover how to do that. What i am working with here is a Nvidia GTX 1070, just so you know. Preparing the project directory and virtual environment Create a new directory/folder whichever terminology your familiar with and cd to it from a terminal instance. Now pay attention here. Diffusers uses PyTorch https://pytorch.org/ as its backend. PyTorch has several wheel repositories for different CUDA versions , and you need to choose for Pascal GPUs the one compatible with CUDA 12.1 or 12.4. To get the right working libraries, we need to fetch them from the repository 'https://download.pytorch.org/whl/cu124' using the --index-url pip CLI parameter/argument. Now, to the Diffusers Framework. Once the PyTorch things is done, the only thing to watch out for is ensuring transformers is fixed at version 5.5.0 . Why is that? My hunch and days of trying to figure out why my setup broke when i updated my libraries in my virtual environment , parts of the transformers API was changed in later version, and diffusers didn't take that into account in their newer versions. From my findings, version 5.5.0 was the last working one. Since we want this to just work, make sure that you only get THAT version until diffusers decides to fix this. Now, lets get to what to run inside our project directory. python3 -m venv ./venv/ ./venv/bin/pip install torch torchvision --index-url 'https://download.pytorch.org/whl/cu124' ./venv/bin/pip install diffusers transformers==5.5.0 accelerate peft This should be all that is needed to get a working virtual environment setup. Using Diffusers for running Stable Diffusion XL models This is now the part in which we need to interact with the Diffusers API. In this case, a regular Python script will do the job, and so, i build one for my needs. python ./venv/bin/python3 import argparse from pathlib import Path import torch from diffusers import DPMSolverMultistepScheduler, StableDiffusionXLPipeline, Constants. ========== IMAGES BATCH COUNT = 2 TORCH DTYPE = torch.bfloat16 Prepare and parse the terminal arguments. ========================================= args parser = argparse.ArgumentParser description="Custom program for text-to-image using Stable Diffusion XL models." args parser.add argument "--model", type=str, required=True, help="Huggingface repository path to the SDXL model weights.", args parser.add argument "--prompt-path", type=Path, default=Path "./prompt.txt" , help="Filepath to the prompt on what to generate.", args parser.add argument "--negative-prompt-path", type=Path, default=None, help="Filepath to the prompt on what NOT to generate. Ignored if not set.", args parser.add argument "--sampling-steps", type=int, default=26, help="How many steps are performed in the denoising procedure. More usually result in better image quality.", args parser.add argument "--guidance-scale", type=float, default=6.5, help="How close it should follow the prompt. Lower values result in more freedom, while higher ones result in more consistency.", args parser.add argument "--images-count", type=int, default=IMAGES BATCH COUNT, help="How many images should be generated per prompt.", args parser.add argument "--seed", type=int, default=None, help="The seed used to set the state of the PRNG. If not set, chooses a random one.", args parser.add argument "--output-dirpath", type=Path, default=Path "./txt2img batch" , help="Where to store the generated images.", parsed args = args parser.parse args Prepare the text-to-image pipeline to generate images. ====================================================== prompt = parsed args.prompt path.read text negative prompt = None if parsed args.negative prompt path == None else parsed args.negative prompt path.read text pipeline = StableDiffusionXLPipeline.from single file parsed args.model, custom pipeline="lpw stable diffusion xl", torch dtype=TORCH DTYPE, variant="bfp16", safety checker=None, pipeline.scheduler = DPMSolverMultistepScheduler.from config pipeline.scheduler.config, use karras sigmas=True, algorithm type="dpmsolver++", pipeline.enable model cpu offload pipeline.enable attention slicing slice size="auto" pipeline.vae.enable slicing Generate and save the images. ============================= batch count, leftover count = divmod parsed args.images count, IMAGES BATCH COUNT images batch sequence = if batch count 0: images batch sequence += IMAGES BATCH COUNT batch count if leftover count 0: images batch sequence += leftover count generator = None if parsed args.seed == None else torch.manual seed parsed args.seed parsed args.output dirpath.mkdir parents=True, exist ok=True, for num images per prompt in images batch sequence: inference result = pipeline prompt=prompt, negative prompt=negative prompt, num inference steps=parsed args.sampling steps, guidance scale=parsed args.guidance scale, num images per prompt=num images per prompt, generator=generator, file counter = 0 saved image paths = for generated image in inference result.images: output filepath = parsed args.output dirpath.joinpath f"image{file counter}.png" while output filepath.exists : file counter += 1 output filepath = parsed args.output dirpath.joinpath f"image{file counter}.png" generated image.save output filepath saved image paths.append output filepath.name print f"Saved the generated {num images per prompt} image/s {saved image paths} into the directory '{parsed args.output dirpath}'." This is what i use to generate AI images with Stable Diffusion XL models. It has a solid CLI interface, and the script can be adapted for ones individual needs if something is insufficient. One important thing to note, due to the Nvidia GTX 1070 being limited to 8 GB of VRAM, and most Stable Diffusion XL models being around 7 GB, reducing the VRAM usage is a must, and this is done usually with offloading https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory offloading , attention slicing https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory memory-efficient-attention and VAE slicing https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory vae-slicing . The best case for offloading is using model offloading https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory model-offloading , but you need to ensure that almost nothing is filling the GPUs VRAM , otherwise the script will likely exit with a out of memory exception when saving the generated images . If that is happening to you way to often, and you have closed all programs, change model offloading to cpu offloading https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory cpu-offloading . It will take longer, but crashes should happen rarely. Even better speed and reliability wise is using group offloading https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory group-offloading with CUDA streaming enabled, but this required having more RAM 16 GB is not enough, 24+ GB is recommended . Regarding attention slicing, it kinda works on the Nvidia GTX 1070, but is very important to enable , since it lowers the VRAM use further, with minimal speed loss. Now, why it is kinda working? Due to FlashAttention https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/fp16 scaled-dot-product-attention the enabled default on newer PyTorch builds simply not working with my GPU, and a altenative being urgently needed. My only option after hours of research and testing ended up being the enable attention slicing https://huggingface.co/docs/diffusers/v0.40.0/en/api/pipelines/overview diffusers.DiffusionPipeline.enable attention slicing method from the base class DiffusionPipeline https://huggingface.co/docs/diffusers/v0.40.0/en/api/pipelines/overview diffusers.DiffusionPipeline class, which does almost the same thing luckily, so we are not totally fucked yet. VAE slicing saves memory by splitting large batches of inputs into a single batch of data and separately processes them. This method works best when generating more than one image at a time, which will be the case most of the time, but keep in mind that inference will take longer the more images you generate. Generating images faster using Ays Now this is mostly uncharted territory, but i was researching how to speed up the inference of Stable Diffusion XL models on my GPU, and found this gem https://research.nvidia.com/labs/toronto-ai/AlignYourSteps/ . Now, how useful is using the AYS schedule for inference? In my experience, it is so-so. Yes, its much faster and does generate good enough images. But...the image quality in regards to prompt adherence is worse, and this depends alot on the model your using. If it was fine tuned very well, it will perform mostly fine. If not, the images will most likely look weird. Also from my findings , it is best to use the sde-dpmsolver++ scheduler https://huggingface.co/docs/diffusers/v0.40.0/en/api/schedulers/multistep dpm solver dpmsolvermultistepscheduler , due to it converging faster and giving good quality results, in 10 steps of inference. Probably best used for experimenting with new ideas. Here is the mostly same script, but using AYS for the inference. python ./venv/bin/python3 import argparse from pathlib import Path import torch from diffusers import DPMSolverMultistepScheduler, StableDiffusionXLPipeline, from diffusers.schedulers import AysSchedules Constants. ========== IMAGES BATCH COUNT = 2 TORCH DTYPE = torch.bfloat16 Prepare and parse the terminal arguments. ========================================= args parser = argparse.ArgumentParser description="Custom program for text-to-image using normal Stable Diffusion models." args parser.add argument "--model", type=str, required=True, help="Huggingface repository path to the SDXL model weights.", args parser.add argument "--prompt-path", type=Path, default=Path "./prompt.txt" , help="Filepath to the prompt on what to generate.", args parser.add argument "--negative-prompt-path", type=Path, default=None, help="Filepath to the prompt on what NOT to generate. Ignored if not set.", args parser.add argument "--guidance-scale", type=float, default=3.0, help="How close it should follow the prompt. Lower values result in more freedom, while higher ones result in more consistency.", args parser.add argument "--images-count", type=int, default=IMAGES BATCH COUNT, help="How many images should be generated per prompt.", args parser.add argument "--seed", type=int, default=None, help="The seed used to set the state of the PRNG. If not set, chooses a random one.", args parser.add argument "--output-dirpath", type=Path, default=Path "./txt2img batch" , help="Where to store the generated images.", parsed args = args parser.parse args Prepare the text-to-image pipeline to generate images. ====================================================== prompt = parsed args.prompt path.read text negative prompt = None if parsed args.negative prompt path == None else parsed args.negative prompt path.read text pipeline = StableDiffusionXLPipeline.from single file parsed args.model, custom pipeline="lpw stable diffusion xl", torch dtype=TORCH DTYPE, safety checker=None, pipeline.scheduler = DPMSolverMultistepScheduler.from config pipeline.scheduler.config, algorithm type="sde-dpmsolver++", solver order=2, pipeline.enable model cpu offload pipeline.enable attention slicing slice size="auto" pipeline.vae.enable slicing Generate and save the images. ============================= batch count, leftover count = divmod parsed args.images count, IMAGES BATCH COUNT images batch sequence = if batch count 0: images batch sequence += IMAGES BATCH COUNT batch count if leftover count 0: images batch sequence += leftover count generator = None if parsed args.seed == None else torch.manual seed parsed args.seed parsed args.output dirpath.mkdir parents=True, exist ok=True, for num images per prompt in images batch sequence: inference result = pipeline prompt=prompt, negative prompt=negative prompt, guidance scale=parsed args.guidance scale, num images per prompt=num images per prompt, generator=generator, timesteps=AysSchedules "StableDiffusionXLTimesteps" , file counter = 0 for generated image in inference result.images: output filepath = parsed args.output dirpath.joinpath f"image{file counter}.png" while output filepath.exists : file counter += 1 output filepath = parsed args.output dirpath.joinpath f"image{file counter}.png" generated image.save output filepath print f"Saved the generated {num images per prompt} image/s into the directory '{parsed args.output dirpath}'." Conclusion It's doable, but will get harder with time, due to lacking software support, and whatever currently works will likely no longer be available in the near future, unless you build the libraries yourself and even that is sometimes a PITA . How do i know that? Why am i this opinion? cuML https://github.com/NVIDIA/cuml and cuDF https://github.com/NVIDIA/cudf . Cannot get binary wheels that have Pascal support, and building them locally failed with cryptic error messages maybe my setup sucks, but it is what it is at the moment . Best i got running is cupy https://github.com/cupy/cupy/ for GPGPU things. So, if you wanna do it, you still can, but do not expect miracles, and inference on the Nvidia 1000 series is not the best the GTX 1070 needs around 3:30 minutes for generating 2 images with model offloading enabled, to give you a idea .