Wiring iOS CoreML to a Quantized On-Device Diffusion Model for Real-Time Image Editing A developer detailed a production pattern for running a quantized Stable Diffusion model on-device in iOS via CoreML, using INT4 weight palettization with FP16 attention projections to cut U-Net size to roughly 750 MB and per-step latency to about 295 ms on an A17 chip. The approach reuses cross-attention key/value caches across denoising steps, cutting cross-attention compute by 35–45% on a 20-step schedule, and applies per-chip memory-tier fallbacks for A16, A17, and M-series devices to avoid silent CPU delegation. --- title: "Wiring CoreML to a Quantized Diffusion Model for Real-Time iOS Editing" published: true description: "Ship quantized Stable Diffusion on iOS with CoreML ML Program format, cross-attention KV-cache reuse, and per-chip memory-tier fallback for A16, A17, and M-series." tags: ios, swift, mobile, architecture canonical url: https://mvpfactory.co/blog/coreml-quantized-diffusion-ios --- Let me show you a pattern I use in every on-device ML project: treat memory pressure as a first-class constraint from day one, not after your first TestFlight crash. Shipping a real-time diffusion-based image editor on iOS is achievable. The UI canvas runs at 60fps because inference executes asynchronously off the main thread — actual per-step latency runs from 280ms to 800ms depending on chip tier and quantization depth. The gap between smooth and jittery comes down to three decisions: quantization depth, attention KV-cache reuse, and a hard per-chip memory ceiling that triggers quality fallback before the OS kills your process. A CoreML inference pipeline for a quantized Stable Diffusion model that: .mlpackage segments coremltools 7.x installed in your Python environment MLModel and MLModelConfiguration A full FP32 SD 1.5 U-Net is unusable on-device. Here is the precision table that matters: | Precision | U-Net Size | ANE Eligible | Step Latency A17 | |---|---|---|---| | FP16 | ~2.5 GB | Partial | ~800 ms/step | | INT8 weights | ~1.3 GB | Yes | ~420 ms/step | | INT4 weights | ~700 MB | Yes | ~280 ms/step | | INT4 + attention FP16 | ~750 MB | Yes | ~295 ms/step | The last row is the production choice. Keep attention projections at FP16 — palettizing them to INT4 compounds error across 20 denoising steps in ways that are visually obvious. Palettize everything else with coremltools.optimize.coreml.palettize weights . This is the highest-leverage optimization most iOS ML engineers skip. In a guided diffusion edit, text conditioning does not change between steps. That means cross-attention K and V projections are identical on every step — recomputing them is pure waste. Here is the minimal setup to get this working: js var cachedKV: String: MLMultiArray = : func denoisingStep latent: MLMultiArray, step: Int throws - MLMultiArray { var inputDict: String: Any = "latent input": latent, "timestep": MLMultiArray step , "use cached kv": MLMultiArray step 0 ? 1 : 0 if step 0 { for key, value in cachedKV { inputDict key = value } } let provider = try MLDictionaryFeatureProvider dictionary: inputDict let output = try unet.prediction from: provider if step == 0 { cachedKV = extractKV from: output } guard let result = output.featureValue for: "latent output" ?.multiArrayValue else { throw InferenceError.missingOutput } return result } This alone cuts cross-attention compute by 35–45% on a 20-step schedule with no quality cost. The docs do not mention this clearly, but MLModelConfiguration.computeUnits should reflect the chip tier detected at runtime. iOS will not crash your app when you breach the Neural Engine's working-set limit — it silently delegates layers to CPU, which is 4–8x slower. | Chip | ANE Budget | Safe Model Budget | Fallback Trigger | |---|---|---|---| | A16 Bionic | ~1.0 GB | ~700 MB | CPU delegation above ~1.1 GB | | A17 Pro | ~1.4 GB | ~1.0 GB | CPU delegation above ~1.5 GB | | M2 / M4 iPad | ~3.5 GB | ~2.5 GB | Rarely triggered | php func resolvedComputeUnits - MLComputeUnits { let chip = ChipTierDetector.current // wrapper around sysctlbyname "hw.optional. " switch chip { case .a16, .a17: return .cpuAndNeuralEngine case .m2, .m4: return .all // GPU path enabled for non-ANE ops default: return .cpuAndNeuralEngine } } Before each inference pass, call os proc available memory . If headroom drops below your model's activation footprint, drop to 384×384 instead of 512×512 rather than letting the runtime decide for you. Silent CPU delegation is your real enemy. There is no error thrown — just latency doubling and frames dropping. Instrument memory headroom before every pass. Benchmarking on M2 iPad and shipping to A16 iPhones. The Neural Engine tier gap is brutal — not just in raw TOPS but in on-chip SRAM for intermediate activations. Always test on your lowest supported chip. Palettizing attention projections. INT4 attention layers compound error visually across 20+ steps. Exempt them explicitly in your palettize weights config and keep them at FP16. Three decisions determine whether your diffusion app ships or gets shelved: INT4 weight palettization with FP16 attention, K/V cache reuse from step zero, and per-chip computeUnits configuration backed by runtime memory monitoring. Each one independently improves the pipeline; together they are what separates a 280ms interactive editor from an 800ms thermally-throttled demo. Resources: