cd /news/ai-tools/wiring-ios-coreml-to-a-quantized-on-… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-142373] src=dev.to β†— pub= topic=ai-tools verified=true sentiment=↑ positive

Wiring iOS CoreML to a Quantized On-Device Diffusion Model for Real-Time Image Editing

A developer detailed a production pattern for running a quantized Stable Diffusion model on-device in iOS via CoreML, using INT4 weight palettization with FP16 attention projections to cut U-Net size to roughly 750 MB and per-step latency to about 295 ms on an A17 chip. The approach reuses cross-attention key/value caches across denoising steps, cutting cross-attention compute by 35–45% on a 20-step schedule, and applies per-chip memory-tier fallbacks for A16, A17, and M-series devices to avoid silent CPU delegation.

by read4 min views1 publishedSep 30, 2026
---
title: "Wiring CoreML to a Quantized Diffusion Model for Real-Time iOS Editing"
published: true
description: "Ship quantized Stable Diffusion on iOS with CoreML ML Program format, cross-attention KV-cache reuse, and per-chip memory-tier fallback for A16, A17, and M-series."
tags: ios, swift, mobile, architecture
canonical_url: https://mvpfactory.co/blog/coreml-quantized-diffusion-ios
---

Let me show you a pattern I use in every on-device ML project: treat memory pressure as a first-class constraint from day one, not after your first TestFlight crash.

Shipping a real-time diffusion-based image editor on iOS is achievable. The UI canvas runs at 60fps because inference executes asynchronously off the main thread β€” actual per-step latency runs from 280ms to 800ms depending on chip tier and quantization depth. The gap between smooth and jittery comes down to three decisions: quantization depth, attention KV-cache reuse, and a hard per-chip memory ceiling that triggers quality fallback before the OS kills your process.

A CoreML inference pipeline for a quantized Stable Diffusion model that:

.mlpackage segmentscoremltools 7.x installed in your Python environmentMLModel and MLModelConfiguration A full FP32 SD 1.5 U-Net is unusable on-device. Here is the precision table that matters:

Precision U-Net Size ANE Eligible Step Latency (A17)
FP16 ~2.5 GB Partial ~800 ms/step
INT8 weights ~1.3 GB Yes ~420 ms/step
INT4 weights ~700 MB Yes ~280 ms/step
INT4 + attention FP16 ~750 MB Yes ~295 ms/step

The last row is the production choice. Keep attention projections at FP16 β€” palettizing them to INT4 compounds error across 20 denoising steps in ways that are visually obvious. Palettize everything else with coremltools.optimize.coreml.palettize_weights.

This is the highest-leverage optimization most iOS ML engineers skip. In a guided diffusion edit, text conditioning does not change between steps. That means cross-attention K and V projections are identical on every step β€” recomputing them is pure waste.

Here is the minimal setup to get this working:

var cachedKV: [String: MLMultiArray] = [:]

func denoisingStep(latent: MLMultiArray, step: Int) throws -> MLMultiArray {
    var inputDict: [String: Any] = [
        "latent_input": latent,
        "timestep": MLMultiArray([step]),
        "use_cached_kv": MLMultiArray([step > 0 ? 1 : 0])
    ]

    if step > 0 {
        for (key, value) in cachedKV { inputDict[key] = value }
    }

    let provider = try MLDictionaryFeatureProvider(dictionary: inputDict)
    let output = try unet.prediction(from: provider)

    if step == 0 {
        cachedKV = extractKV(from: output)
    }

    guard let result = output.featureValue(for: "latent_output")?.multiArrayValue else {
        throw InferenceError.missingOutput
    }
    return result
}

This alone cuts cross-attention compute by 35–45% on a 20-step schedule with no quality cost.

The docs do not mention this clearly, but MLModelConfiguration.computeUnits should reflect the chip tier detected at runtime. iOS will not crash your app when you breach the Neural Engine's working-set limit β€” it silently delegates layers to CPU, which is 4–8x slower.

Chip ANE Budget Safe Model Budget Fallback Trigger
A16 Bionic ~1.0 GB ~700 MB CPU delegation above ~1.1 GB
A17 Pro ~1.4 GB ~1.0 GB CPU delegation above ~1.5 GB
M2 / M4 (iPad) ~3.5 GB ~2.5 GB Rarely triggered
func resolvedComputeUnits() -> MLComputeUnits {
    let chip = ChipTierDetector.current() // wrapper around sysctlbyname("hw.optional.*")
    switch chip {
    case .a16, .a17:
        return .cpuAndNeuralEngine
    case .m2, .m4:
        return .all // GPU path enabled for non-ANE ops
    default:
        return .cpuAndNeuralEngine
    }
}

Before each inference pass, call os_proc_available_memory(). If headroom drops below your model's activation footprint, drop to 384Γ—384 instead of 512Γ—512 rather than letting the runtime decide for you.

Silent CPU delegation is your real enemy. There is no error thrown β€” just latency doubling and frames dropping. Instrument memory headroom before every pass.

Benchmarking on M2 iPad and shipping to A16 iPhones. The Neural Engine tier gap is brutal β€” not just in raw TOPS but in on-chip SRAM for intermediate activations. Always test on your lowest supported chip.

Palettizing attention projections. INT4 attention layers compound error visually across 20+ steps. Exempt them explicitly in your palettize_weights config and keep them at FP16.

Three decisions determine whether your diffusion app ships or gets shelved: INT4 weight palettization with FP16 attention, K/V cache reuse from step zero, and per-chip computeUnits configuration backed by runtime memory monitoring. Each one independently improves the pipeline; together they are what separates a 280ms interactive editor from an 800ms thermally-throttled demo.

Resources:

── more in #ai-tools 4 stories Β· sorted by recency
── more on @coreml 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/wiring-ios-coreml-to…] indexed:0 read:4min 2026-09-30 Β· β€”