Qwen-Image-2.1 runs in diffusers through QwenImage21Pipeline. In bfloat16 its text encoder takes 16.3 GB, its transformer 13.3 GB and its VAE 1.3 GB, so the loaded pipeline came to 31.4 GB on a 48 GB Mac, and decoding a 1024 × 1024 image pushed it 11 GB higher. This guide shows what the usual setup did on that Mac, why model CPU offload did not help, and the setup that finished: one of the two large models in memory at a time.
The runs used an Apple M5 Pro with 48 GB, torch 2.14.0 and diffusers 0.41.0.dev0 from main. QwenImage21Pipeline also ships in the diffusers 0.41.0 release, which I did not run. Every run made one 1024 × 1024 image in 20 steps with seed 7, under a memory guard that stopped the run once swap had grown by 8 GB.
The usual code loads all three models onto the GPU and keeps them there:
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("mps")
image = pipe(prompt=prompt, width=1024, height=1024, num_inference_steps=20).images[0]
It got through encoding and all 20 denoising steps at 31 to 33 GB. The VAE decode then took the footprint from 32.5 to 43.6 GB and swap grew by 8 GB, so the guard stopped the run before it wrote an image. PyTorch raised no out-of-memory error before that; the Mac went into swap instead.
The answer diffusers has built in for a pipeline that does not fit is model CPU offload: each model waits on the CPU and moves to the GPU only while it runs.
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload(device="mps")
On a Mac the CPU and the GPU share one pool of memory, so a model that waits "on the CPU" still sits in the same RAM. In the one offload run, the process footprint never went above 18.5 GB, which looks like a fit. Swap grew anyway: by 3.5 GB during denoising and to 9.6 GB in the VAE decode, and the guard stopped the run where it had stopped eager. The footprint, the number Activity Monitor shows for a process, does not count pages that are already in swap, so it cannot show this.
The pipeline uses its models in turn: the text encoder only to encode the prompt, the transformer only in the denoising loop, and neither in the VAE decode. So the two large models never have to be in memory together. stageload keeps track of that: StagedModels loads a model the first time a stage asks for it and releases it when the pipeline enters a stage that does not list it.
from stageload import StagedModels, on_call
STAGES = {"encode": ["text_encoder"], "denoise": ["transformer"], "decode": []}
models = StagedModels({"text_encoder": text_encoder, "transformer": transformer}, stages=STAGES)
on_call(pipe, "encode_prompt", before=lambda *a, **k: models.enter("encode"))
on_call(pipe, "prepare_latents", before=lambda *a, **k: models.enter("denoise"))
on_call(pipe, "_unpack_latents", before=lambda *a, **k: models.enter("decode"))
text_encoder and transformer are functions that load each model with from_pretrained(..., subfolder=...) and move it to mps. The pipeline itself is built with text_encoder=None, transformer=None and the VAE, which stays loaded. on_call runs models.enter(stage) just before the pipeline methods where each stage begins, so the pipeline's code stays as it is. One more piece is needed: the pipeline reads the transformer's config before denoising starts, so pipe.text_encoder and pipe.transformer are stand-ins that answer .config from the checkpoint and load the model on any other use. The whole script, with the stand-ins and a memory trace, is bench/qwen_image.py, and the library installs with pip install stageload.
Both staged runs finished, in 94 and 124 s. The footprint peaked at 19.0 GB while encoding, at 16.2 and 18.6 GB in the two runs' denoising, and at 14.4 GB in the decode. Swap did not grow in either run, and free memory never fell below 51 %. the two models took about 13 s per run, and releasing the transformer before the decode another 4 to 5 s. Counted from the moment the transformer was in memory to the start of the decode, the 20 steps took 70 s in one staged run, 99 s in the other and 77.5 s for eager; the traces do not show why the two staged runs differ. Both produced the same image, bit for bit.
A run fits when swap does not grow while it runs, whatever the footprint says. Two commands show it:
sysctl vm.swapusage # compare "used" before and after the run
sysctl kern.memorystatus_level # system-wide free memory, in percent
stageload's MemoryMeter writes the footprint, swap and free memory to a JSONL trace every 0.5 s, and stageload guard starts a command when there is room for it and stops it once swap has grown past a budget:
stageload guard --wait-free 40 --swap-budget 8G -- python generate.py
pipe.vae.enable_tiling() decodes the image in tiles, and the decode is where both eager and offload went into swap. I did not try it, so I cannot say whether it lets eager finish on this Mac.
The measurements in more detail are in Qwen-Image-2.1 one stage at a time and Model CPU offload on a 48 GB Mac, and the traces are in the stageload repository.