# Wiring iOS CoreML to a Quantized On-Device Diffusion Model for Real-Time Image Editing

> Source: <https://dev.to/software_mvp-factory/wiring-ios-coreml-to-a-quantized-on-device-diffusion-model-for-real-time-image-editing-39l8>
> Published: 2026-09-30 08:39:09+00:00



```
---
title: "Wiring CoreML to a Quantized Diffusion Model for Real-Time iOS Editing"
published: true
description: "Ship quantized Stable Diffusion on iOS with CoreML ML Program format, cross-attention KV-cache reuse, and per-chip memory-tier fallback for A16, A17, and M-series."
tags: ios, swift, mobile, architecture
canonical_url: https://mvpfactory.co/blog/coreml-quantized-diffusion-ios
---
```

Let me show you a pattern I use in every on-device ML project: treat memory pressure as a first-class constraint from day one, not after your first TestFlight crash.

Shipping a real-time diffusion-based image editor on iOS is achievable. The UI canvas runs at 60fps because inference executes asynchronously off the main thread — actual per-step latency runs from 280ms to 800ms depending on chip tier and quantization depth. The gap between smooth and jittery comes down to three decisions: quantization depth, attention KV-cache reuse, and a hard per-chip memory ceiling that triggers quality fallback before the OS kills your process.

A CoreML inference pipeline for a quantized Stable Diffusion model that:

`.mlpackage` segments`coremltools` 7.x installed in your Python environment`MLModel` and `MLModelConfiguration`
A full FP32 SD 1.5 U-Net is unusable on-device. Here is the precision table that matters:

| Precision | U-Net Size | ANE Eligible | Step Latency (A17) | 
|---|---|---|---|
| FP16 | ~2.5 GB | Partial | ~800 ms/step | 
| INT8 weights | ~1.3 GB | Yes | ~420 ms/step | 
| INT4 weights | ~700 MB | Yes | ~280 ms/step | 
| INT4 + attention FP16 | ~750 MB | Yes | ~295 ms/step | 

The last row is the production choice. Keep attention projections at FP16 — palettizing them to INT4 compounds error across 20 denoising steps in ways that are visually obvious. Palettize everything else with `coremltools.optimize.coreml.palettize_weights`.

This is the highest-leverage optimization most iOS ML engineers skip. In a guided diffusion edit, text conditioning does not change between steps. That means cross-attention K and V projections are identical on every step — recomputing them is pure waste.

Here is the minimal setup to get this working:

``` js
var cachedKV: [String: MLMultiArray] = [:]

func denoisingStep(latent: MLMultiArray, step: Int) throws -> MLMultiArray {
    var inputDict: [String: Any] = [
        "latent_input": latent,
        "timestep": MLMultiArray([step]),
        "use_cached_kv": MLMultiArray([step > 0 ? 1 : 0])
    ]

    if step > 0 {
        for (key, value) in cachedKV { inputDict[key] = value }
    }

    let provider = try MLDictionaryFeatureProvider(dictionary: inputDict)
    let output = try unet.prediction(from: provider)

    if step == 0 {
        cachedKV = extractKV(from: output)
    }

    guard let result = output.featureValue(for: "latent_output")?.multiArrayValue else {
        throw InferenceError.missingOutput
    }
    return result
}
```

This alone cuts cross-attention compute by 35–45% on a 20-step schedule with no quality cost.

The docs do not mention this clearly, but `MLModelConfiguration.computeUnits` should reflect the chip tier detected at runtime. iOS will not crash your app when you breach the Neural Engine's working-set limit — it silently delegates layers to CPU, which is 4–8x slower.

| Chip | ANE Budget | Safe Model Budget | Fallback Trigger | 
|---|---|---|---|
| A16 Bionic | ~1.0 GB | ~700 MB | CPU delegation above ~1.1 GB | 
| A17 Pro | ~1.4 GB | ~1.0 GB | CPU delegation above ~1.5 GB | 
| M2 / M4 (iPad) | ~3.5 GB | ~2.5 GB | Rarely triggered | 

``` php
func resolvedComputeUnits() -> MLComputeUnits {
    let chip = ChipTierDetector.current() // wrapper around sysctlbyname("hw.optional.*")
    switch chip {
    case .a16, .a17:
        return .cpuAndNeuralEngine
    case .m2, .m4:
        return .all // GPU path enabled for non-ANE ops
    default:
        return .cpuAndNeuralEngine
    }
}
```

Before each inference pass, call `os_proc_available_memory()`. If headroom drops below your model's activation footprint, drop to 384×384 instead of 512×512 rather than letting the runtime decide for you.

**Silent CPU delegation is your real enemy.** There is no error thrown — just latency doubling and frames dropping. Instrument memory headroom before every pass.

**Benchmarking on M2 iPad and shipping to A16 iPhones.** The Neural Engine tier gap is brutal — not just in raw TOPS but in on-chip SRAM for intermediate activations. Always test on your lowest supported chip.

**Palettizing attention projections.** INT4 attention layers compound error visually across 20+ steps. Exempt them explicitly in your `palettize_weights` config and keep them at FP16.

Three decisions determine whether your diffusion app ships or gets shelved: INT4 weight palettization with FP16 attention, K/V cache reuse from step zero, and per-chip `computeUnits` configuration backed by runtime memory monitoring. Each one independently improves the pipeline; together they are what separates a 280ms interactive editor from an 800ms thermally-throttled demo.

**Resources:**
