{"slug": "setup-diffusers-for-stable-diffusion-xl-using-a-nvidia-pascal-gpu-in-linux", "title": "Setup Diffusers for Stable Diffusion XL Using a Nvidia Pascal GPU in Linux", "summary": "A developer documented a working Linux setup for running Stable Diffusion XL through Hugging Face Diffusers on an Nvidia GTX 1070 Pascal GPU, pinning PyTorch to the CUDA 12.4 wheel index and transformers to version 5.5.0. The guide states that 5.5.0 was the last transformers release compatible with newer Diffusers versions, and provides a Python script using StableDiffusionXLPipeline with DPMSolverMultistepScheduler, 26 sampling steps, a 6.5 guidance scale, and bfloat16 precision to generate two images per prompt.", "body_md": "You need to first ensure that all the necessary drivers (CUDA most important) for the Nvidia GPU are installed. This post will NOT cover how to do that.\n\nWhat i am working with here is a Nvidia GTX 1070, just so you know.\n\n## Preparing the project directory and virtual environment\n\nCreate a new directory/folder (whichever terminology your familiar with) and `cd` to it from a terminal instance.\n\n**Now pay attention here.**\n\nDiffusers uses [PyTorch](https://pytorch.org/) as its backend.\nPyTorch has *several wheel repositories for different CUDA versions*, and you need to choose for Pascal GPUs the one compatible with CUDA 12.1 or 12.4.\n\nTo get the right working libraries, we need to fetch them from the repository 'https://download.pytorch.org/whl/cu124' using the `--index-url` pip CLI parameter/argument.\n\nNow, to the Diffusers Framework.\n\nOnce the PyTorch things is done,\nthe only thing to watch out for is ensuring `transformers` is fixed at version `5.5.0` .\nWhy is that? My hunch (and days of trying to figure out why my setup broke when i updated my libraries in my virtual environment), \nparts of the `transformers` API was changed in later version, and `diffusers` didn't take that into account in their newer versions.\nFrom my findings, version `5.5.0` was the last working one. Since we want this to just work, \nmake sure that you only get THAT version until `diffusers` decides to fix this.\n\nNow, lets get to what to run inside our project directory.\n\n```\npython3 -m venv ./venv/\n./venv/bin/pip install torch torchvision --index-url 'https://download.pytorch.org/whl/cu124'\n./venv/bin/pip install diffusers transformers==5.5.0 accelerate peft\n```\n\nThis should be all that is needed to get a working virtual environment setup.\n\n## Using Diffusers for running Stable Diffusion XL models\n\nThis is now the part in which we need to interact with the Diffusers API. In this case, a regular Python script will do the job, and so, i build one for my needs.\n\n``` python\n#!./venv/bin/python3\nimport argparse\nfrom pathlib import Path\n\nimport torch\nfrom diffusers import (\n    DPMSolverMultistepScheduler,\n    StableDiffusionXLPipeline,\n)\n\n# Constants.\n# ==========\nIMAGES_BATCH_COUNT = 2\nTORCH_DTYPE = torch.bfloat16\n\n# Prepare and parse the terminal arguments.\n# =========================================\nargs_parser = argparse.ArgumentParser(\n    description=\"Custom program for text-to-image using Stable Diffusion XL models.\"\n)\nargs_parser.add_argument(\n    \"--model\",\n    type=str,\n    required=True,\n    help=\"Huggingface repository path to the SDXL model weights.\",\n)\nargs_parser.add_argument(\n    \"--prompt-path\",\n    type=Path,\n    default=Path(\"./prompt.txt\"),\n    help=\"Filepath to the prompt on what to generate.\",\n)\nargs_parser.add_argument(\n    \"--negative-prompt-path\",\n    type=Path,\n    default=None,\n    help=\"Filepath to the prompt on what NOT to generate. Ignored if not set.\",\n)\nargs_parser.add_argument(\n    \"--sampling-steps\",\n    type=int,\n    default=26,\n    help=\"How many steps are performed in the denoising procedure. More usually result in better image quality.\",\n)\nargs_parser.add_argument(\n    \"--guidance-scale\",\n    type=float,\n    default=6.5,\n    help=\"How close it should follow the prompt. Lower values result in more freedom, while higher ones result in more consistency.\",\n)\nargs_parser.add_argument(\n    \"--images-count\",\n    type=int,\n    default=IMAGES_BATCH_COUNT,\n    help=\"How many images should be generated per prompt.\",\n)\nargs_parser.add_argument(\n    \"--seed\",\n    type=int,\n    default=None,\n    help=\"The seed used to set the state of the PRNG. If not set, chooses a random one.\",\n)\nargs_parser.add_argument(\n    \"--output-dirpath\",\n    type=Path,\n    default=Path(\"./txt2img_batch\"),\n    help=\"Where to store the generated images.\",\n)\n\nparsed_args = args_parser.parse_args()\n\n# Prepare the text-to-image pipeline to generate images.\n# ======================================================\nprompt = parsed_args.prompt_path.read_text()\nnegative_prompt = None if parsed_args.negative_prompt_path == None else parsed_args.negative_prompt_path.read_text()\n\npipeline = StableDiffusionXLPipeline.from_single_file(\n    parsed_args.model,\n    custom_pipeline=\"lpw_stable_diffusion_xl\",\n    torch_dtype=TORCH_DTYPE,\n    variant=\"bfp16\",\n    safety_checker=None,\n)\npipeline.scheduler = DPMSolverMultistepScheduler.from_config(\n    pipeline.scheduler.config,\n    use_karras_sigmas=True,\n    algorithm_type=\"dpmsolver++\",\n)\npipeline.enable_model_cpu_offload()\npipeline.enable_attention_slicing(slice_size=\"auto\")\npipeline.vae.enable_slicing()\n\n# Generate and save the images.\n# =============================\nbatch_count, leftover_count = divmod(parsed_args.images_count, IMAGES_BATCH_COUNT)\nimages_batch_sequence = []\n\nif batch_count > 0:\n    images_batch_sequence += [IMAGES_BATCH_COUNT] * batch_count\nif leftover_count > 0:\n    images_batch_sequence += [leftover_count]\n\ngenerator = None if parsed_args.seed == None else torch.manual_seed(parsed_args.seed)\n\nparsed_args.output_dirpath.mkdir(\n    parents=True,\n    exist_ok=True,\n)\nfor num_images_per_prompt in images_batch_sequence:\n    inference_result = pipeline(\n        prompt=prompt,\n        negative_prompt=negative_prompt,\n        num_inference_steps=parsed_args.sampling_steps,\n        guidance_scale=parsed_args.guidance_scale,\n        num_images_per_prompt=num_images_per_prompt,\n        generator=generator,\n    )\n\n    file_counter = 0\n    saved_image_paths = []\n    for generated_image in inference_result.images:\n        output_filepath = parsed_args.output_dirpath.joinpath(f\"image{file_counter}.png\")\n        while output_filepath.exists():\n            file_counter += 1\n            output_filepath = parsed_args.output_dirpath.joinpath(f\"image{file_counter}.png\")\n\n        generated_image.save(output_filepath)\n        saved_image_paths.append(output_filepath.name)\n\n    print(f\"Saved the generated {num_images_per_prompt} image/s {saved_image_paths} into the directory '{parsed_args.output_dirpath}'.\")\n```\n\nThis is what i use to generate AI images with Stable Diffusion XL models. It has a solid CLI interface, and the script can be adapted for ones individual needs if something is insufficient.\n\nOne important thing to note, due to the Nvidia GTX 1070 being limited to 8 GB of VRAM, and most Stable Diffusion XL models being around 7 GB,\nreducing the VRAM usage is a must, and this is done usually with [offloading](https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory#offloading), [attention slicing](https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory#memory-efficient-attention) and [VAE slicing](https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory#vae-slicing).\n\nThe best case for offloading is using [model offloading](https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory#model-offloading), but you need to ensure that *almost nothing is filling the GPUs VRAM*, otherwise the script will likely exit with a *out of memory exception when saving the generated images*. If that is happening to you way to often, and you have closed all programs, change model offloading to [cpu offloading](https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory#cpu-offloading). It will take longer, but crashes should happen rarely. Even better (speed and reliability wise) is using [group offloading](https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/memory#group-offloading) with CUDA streaming enabled, but this required having more RAM (16 GB is not enough, 24+ GB is recommended).\n\nRegarding attention slicing, it kinda works on the Nvidia GTX 1070, but is **very important to enable**, since it lowers the VRAM use further, with minimal speed loss.\nNow, why it is kinda working? Due to [FlashAttention](https://huggingface.co/docs/diffusers/v0.40.0/en/optimization/fp16#scaled-dot-product-attention) (the enabled default on newer PyTorch builds) simply not working with my GPU, and a altenative being urgently needed.\nMy only option (after hours of research and testing) ended up being the [enable_attention_slicing](https://huggingface.co/docs/diffusers/v0.40.0/en/api/pipelines/overview#diffusers.DiffusionPipeline.enable_attention_slicing) method from the base class [DiffusionPipeline](https://huggingface.co/docs/diffusers/v0.40.0/en/api/pipelines/overview#diffusers.DiffusionPipeline) class, which does (almost) the same thing luckily, so we are not totally fucked yet.\n\nVAE slicing saves memory by splitting large batches of inputs into a single batch of data and separately processes them. This method works best when generating more than one image at a time, which will be the case most of the time, but keep in mind that inference will take longer the more images you generate.\n\n## Generating images faster using Ays\n\nNow this is mostly uncharted territory, but i was researching how to speed up the inference of Stable Diffusion XL models\non my GPU, and found [this gem](https://research.nvidia.com/labs/toronto-ai/AlignYourSteps/).\n\nNow, how useful is using the AYS schedule for inference? In my experience, it is so-so.\n\nYes, its much faster and does generate good enough images. But...the image quality in regards to prompt adherence is worse, and this depends alot on the model your using. If it was fine tuned very well, it will perform mostly fine. If not, the images will most likely look weird.\n\nAlso (from my findings), it is best to use the [sde-dpmsolver++ scheduler](https://huggingface.co/docs/diffusers/v0.40.0/en/api/schedulers/multistep_dpm_solver#dpmsolvermultistepscheduler), \ndue to it converging faster and giving good quality results, in 10 steps of inference.\n\nProbably best used for experimenting with new ideas.\n\nHere is the (mostly) same script, but using AYS for the inference.\n\n``` python\n#!./venv/bin/python3\nimport argparse\nfrom pathlib import Path\n\nimport torch\nfrom diffusers import (\n    DPMSolverMultistepScheduler,\n    StableDiffusionXLPipeline,\n)\nfrom diffusers.schedulers import AysSchedules\n\n# Constants.\n# ==========\nIMAGES_BATCH_COUNT = 2\nTORCH_DTYPE = torch.bfloat16\n\n# Prepare and parse the terminal arguments.\n# =========================================\nargs_parser = argparse.ArgumentParser(\n    description=\"Custom program for text-to-image using normal Stable Diffusion models.\"\n)\nargs_parser.add_argument(\n    \"--model\",\n    type=str,\n    required=True,\n    help=\"Huggingface repository path to the SDXL model weights.\",\n)\nargs_parser.add_argument(\n    \"--prompt-path\",\n    type=Path,\n    default=Path(\"./prompt.txt\"),\n    help=\"Filepath to the prompt on what to generate.\",\n)\nargs_parser.add_argument(\n    \"--negative-prompt-path\",\n    type=Path,\n    default=None,\n    help=\"Filepath to the prompt on what NOT to generate. Ignored if not set.\",\n)\nargs_parser.add_argument(\n    \"--guidance-scale\",\n    type=float,\n    default=3.0,\n    help=\"How close it should follow the prompt. Lower values result in more freedom, while higher ones result in more consistency.\",\n)\nargs_parser.add_argument(\n    \"--images-count\",\n    type=int,\n    default=IMAGES_BATCH_COUNT,\n    help=\"How many images should be generated per prompt.\",\n)\nargs_parser.add_argument(\n    \"--seed\",\n    type=int,\n    default=None,\n    help=\"The seed used to set the state of the PRNG. If not set, chooses a random one.\",\n)\nargs_parser.add_argument(\n    \"--output-dirpath\",\n    type=Path,\n    default=Path(\"./txt2img_batch\"),\n    help=\"Where to store the generated images.\",\n)\n\nparsed_args = args_parser.parse_args()\n\n# Prepare the text-to-image pipeline to generate images.\n# ======================================================\nprompt = parsed_args.prompt_path.read_text()\nnegative_prompt = None if parsed_args.negative_prompt_path == None else parsed_args.negative_prompt_path.read_text()\n\npipeline = StableDiffusionXLPipeline.from_single_file(\n    parsed_args.model,\n    custom_pipeline=\"lpw_stable_diffusion_xl\",\n    torch_dtype=TORCH_DTYPE,\n    safety_checker=None,\n)\npipeline.scheduler = DPMSolverMultistepScheduler.from_config(\n    pipeline.scheduler.config,\n    algorithm_type=\"sde-dpmsolver++\",\n    solver_order=2,\n)\npipeline.enable_model_cpu_offload()\npipeline.enable_attention_slicing(slice_size=\"auto\")\npipeline.vae.enable_slicing()\n\n# Generate and save the images.\n# =============================\nbatch_count, leftover_count = divmod(parsed_args.images_count, IMAGES_BATCH_COUNT)\nimages_batch_sequence = []\n\nif batch_count > 0:\n    images_batch_sequence += [IMAGES_BATCH_COUNT] * batch_count\nif leftover_count > 0:\n    images_batch_sequence += [leftover_count]\n\ngenerator = None if parsed_args.seed == None else torch.manual_seed(parsed_args.seed)\n\nparsed_args.output_dirpath.mkdir(\n    parents=True,\n    exist_ok=True,\n)\nfor num_images_per_prompt in images_batch_sequence:\n    inference_result = pipeline(\n        prompt=prompt,\n        negative_prompt=negative_prompt,\n        guidance_scale=parsed_args.guidance_scale,\n        num_images_per_prompt=num_images_per_prompt,\n        generator=generator,\n        timesteps=AysSchedules[\"StableDiffusionXLTimesteps\"],\n    )\n\n    file_counter = 0\n    for generated_image in inference_result.images:\n        output_filepath = parsed_args.output_dirpath.joinpath(f\"image{file_counter}.png\")\n        while output_filepath.exists():\n            file_counter += 1\n            output_filepath = parsed_args.output_dirpath.joinpath(f\"image{file_counter}.png\")\n\n        generated_image.save(output_filepath)\n\n    print(f\"Saved the generated {num_images_per_prompt} image/s into the directory '{parsed_args.output_dirpath}'.\")\n```\n\n## Conclusion\n\nIt's doable, but will get harder with time, due to lacking software support, and whatever currently works will likely no longer be available in the near future,\nunless you build the libraries yourself (and even that is sometimes a PITA). How do i know that? Why am i this opinion? [cuML](https://github.com/NVIDIA/cuml) and [cuDF](https://github.com/NVIDIA/cudf). Cannot get binary wheels that have Pascal support,\nand building them locally failed with cryptic error messages (maybe my setup sucks, but it is what it is at the moment).\nBest i got running is [cupy](https://github.com/cupy/cupy/) for GPGPU things.\n\nSo, if you wanna do it, you still can, but do not expect miracles, and inference on the Nvidia 1000 series is not the best (the GTX 1070 needs around 3:30 minutes for generating 2 images with model offloading enabled, to give you a idea).", "url": "https://wpnews.pro/news/setup-diffusers-for-stable-diffusion-xl-using-a-nvidia-pascal-gpu-in-linux", "canonical_source": "https://zansf-personal-blog.statichost.page/how-to-setup-diffusers-for-stable-diffusion-xl-using-a-nvidia-pascal-gpu-in-linux-ubuntu-or-debian-based.html", "published_at": "2026-10-02 13:57:06+00:00", "updated_at": "2026-10-02 14:06:45.265352+00:00", "lang": "en", "topics": ["generative-ai", "ai-tools", "ai-infrastructure", "developer-tools"], "entities": ["Nvidia", "GTX 1070", "Hugging Face", "Diffusers", "Stable Diffusion XL", "PyTorch", "transformers", "CUDA"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/setup-diffusers-for-stable-diffusion-xl-using-a-nvidia-pascal-gpu-in-linux", "markdown": "https://wpnews.pro/news/setup-diffusers-for-stable-diffusion-xl-using-a-nvidia-pascal-gpu-in-linux.md", "text": "https://wpnews.pro/news/setup-diffusers-for-stable-diffusion-xl-using-a-nvidia-pascal-gpu-in-linux.txt", "jsonld": "https://wpnews.pro/news/setup-diffusers-for-stable-diffusion-xl-using-a-nvidia-pascal-gpu-in-linux.jsonld"}}