{"slug": "wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented", "title": "Wiring iOS CoreML to a Quantized On-Device Reranker for Retrieval-Augmented Generation", "summary": "A developer outlined a two-stage on-device retrieval-augmented generation pipeline for iPhone that pairs an INT8-quantized MiniLM-L6 bi-encoder for approximate retrieval with a CoreML cross-encoder (MiniLM-L12) that rescores the top candidates, targeting an 80–120ms end-to-end latency budget on A16/A17 chips. The writeup reports that INT8 quantization costs under 0.5% NDCG on the bi-encoder and about 1.2% on the cross-encoder, while INT4 drops roughly 4.8% NDCG, and that reusing query-side KV states across candidates is a larger latency win than quantization alone.", "body_md": "\n\n```\n---\ntitle: \"CoreML Cross-Encoder Reranking: On-Device RAG on iPhone\"\npublished: true\ndescription: \"Build a two-stage on-device RAG pipeline on iPhone using CoreML — INT8-quantized bi-encoder retrieval, cross-encoder reranking, and A16/A17 latency budgets explained.\"\ntags: ios, swift, mobile, architecture\ncanonical_url: https://mvpfactory.co/blog/coreml-cross-encoder-reranking-on-device-rag-iphone\n---\n\n## What We Are Building\n\nBy the end of this walkthrough, you will have a production-grade two-stage RAG pipeline running entirely on-device. Stage one: fast approximate retrieval with a quantized bi-encoder. Stage two: a CoreML cross-encoder that rescores the top-k candidates. The latency budget on A16/A17 chips is roughly 80–120ms end-to-end before users notice lag. Let me show you the pattern I use to get there.\n\n## Prerequisites\n\n- Xcode 15+\n- `coremltools` 7.x installed in your Python environment\n- A pre-trained bi-encoder (MiniLM-L6) and cross-encoder (MiniLM-L12) in PyTorch\n- Basic familiarity with CoreML model conversion\n\n---\n\n## The Two-Stage Architecture\n\nHere is the mental model worth internalising before writing a single line of code:\n```\n\nQuery → [Bi-Encoder] → embedding → ANN search (top-k=50)\n\n                                          ↓\n\n                              [Cross-Encoder] → rescore top-k=10\n\n                                          ↓\n\n                                  Final ranked results\n\n```\nThe bi-encoder runs **once per query**. The cross-encoder runs **k times**. This asymmetry is the entire reason the two-stage model exists — cross-encoders are dramatically more accurate but cannot scale to full corpus retrieval.\n\n---\n\n## Step 1: Quantization — Get the Numbers Right\n\nMost teams get this wrong: they quantize both stages with the same strategy and wonder why quality collapses. The bi-encoder and cross-encoder have different sensitivity profiles.\n\n| Model Stage | FP32 Size | INT8 Size | Latency (A17) | Quality Drop |\n|---|---|---|---|---|\n| Bi-encoder (MiniLM-L6) | 90 MB | 24 MB | 8 ms | < 0.5% NDCG |\n| Cross-encoder (MiniLM-L12) | 180 MB | 47 MB | 34 ms/candidate | ~1.2% NDCG |\n| Cross-encoder (INT4 aggressive) | 180 MB | 23 MB | 19 ms/candidate | ~4.8% NDCG |\n\nINT8 on the bi-encoder is nearly free. On the cross-encoder, INT8 is acceptable; INT4 starts hurting meaningful recall. Stay at INT8 for cross-encoders in production.\n\nExport with CoreML Tools like this:\n```\n\npython\n\nimport coremltools as ct\n\nmlmodel = ct.convert(\n\n    traced_model,\n\n    compute_precision=ct.precision.FLOAT16,\n\n    compute_units=ct.ComputeUnit.ALL  # uses ANE + GPU\n\n)\n\nmlmodel.save(\"CrossEncoder.mlpackage\")\n\n```\nTarget `ct.ComputeUnit.ALL` to push attention layers onto the Neural Engine. On A16 and A17, this cuts cross-encoder latency by roughly 40% versus CPU-only execution.\n\n---\n\n## Step 2: KV-Cache Reuse Across Candidates\n\nHere is the gotcha that will save you hours. The biggest latency win in cross-encoder scoring is not quantization — it is KV-cache reuse. For a given query, the query-side key/value representations are identical across all k candidates. You compute them once.\n```\n\nswift\n\n// Cache query KV states once per query\n\nlet queryKV = crossEncoder.encodeQuery(queryTokens)\n\nlet scores = candidates.map { doc in\n\n    crossEncoder.scoreWithCachedQuery(queryKV, docTokens: doc.tokens)\n\n}\n\n```\nCoreML does not expose KV-cache injection directly for transformer models. You have two options: split the model at the cross-attention boundary and pass cached states manually, or use ONNX Runtime with a custom CoreML execution provider. Go with the split-model approach. It avoids the ONNX Runtime dependency, stays entirely within the Apple toolchain, and is straightforward to debug with Instruments.\n\n---\n\n## Step 3: Respect the Memory Pressure Ceiling\n\nThe constraint nobody mentions until production: the Neural Engine on A16/A17 shares memory bandwidth with the GPU, CPU, and camera ISP. Under sustained load, the kernel will thermally throttle ANE frequency within 60–90 seconds.\n\nThe practical ceiling:\n\n- **k = 10–20 candidates**: Safe. Stays below 200ms total, no throttle.\n- **k = 30–50 candidates**: Borderline. Latency climbs to 400–600ms. Sustained sessions trigger throttle at ~45s.\n- **k > 50**: Avoid entirely on-device unless you batch overnight with background processing.\n\nFor interactive sessions, monitor thermal state and shed load before the system does it for you:\n```\n\nswift\n\nlet thermal = NSProcessInfo.processInfo.thermalState\n\nif thermal == .serious || thermal == .critical {\n\n    // reduce k or defer reranking\n\n}\n\n```\nFor sustained background workloads, schedule reranking via `BGProcessingTaskRequest`, which the OS grants during charging when thermals are favorable.\n\n---\n\n## Gotchas\n\n**INT4 looks tempting — resist it.** The 2× size reduction is not worth the ~5% NDCG regression on cross-encoders. Validate against your own eval set (200–500 query/document pairs, scored with `pytrec_eval`) before committing to any quantization strategy.\n\n**The working set math matters.** For a 12-layer cross-encoder scoring 20 candidates of 512 tokens, you are looking at ~310MB active — fine. At 50 candidates, you are pushing 750MB and competing with the OS.\n\n**Verify quality regressions before shipping.** A 5-minute offline eval loop with `pytrec_eval` catches regressions between FP32 and quantized outputs. The docs do not mention this, but skipping it is how quality silently degrades in production.\n\nSpeaking of sustained focus work — if you are doing long CoreML optimization sessions, [HealthyDesk](https://play.google.com/store/apps/details?id=com.healthydesk) is worth keeping open in the background. Thermal throttling your chips and your own posture in the same session is a bad combo.\n\n---\n\n## Conclusion\n\nThree things worth remembering:\n\n1. **Use INT8 for both stages, not INT4.** Validate against your eval set before committing.\n2. **Implement query-side KV-cache reuse via the split-model approach.** This cuts reranking latency by 30–45% on A16/A17 with no quality cost.\n3. **Cap candidate count at k=20 for interactive sessions.** Schedule anything heavier via `BGProcessingTaskRequest`.\n\nThe two-stage pipeline is the right architecture for on-device RAG. Get the quantization strategy and the memory ceiling right, and you will stay comfortably inside the 80–120ms budget users never notice.\n```\n\n", "url": "https://wpnews.pro/news/wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented", "canonical_source": "https://dev.to/software_mvp-factory/wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented-generation-1o08", "published_at": "2026-09-29 08:38:42+00:00", "updated_at": "2026-09-29 08:46:44.543995+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "developer-tools", "ai-tools"], "entities": ["CoreML", "Apple", "MiniLM-L6", "MiniLM-L12", "coremltools", "Xcode", "Neural Engine", "ONNX Runtime"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented", "markdown": "https://wpnews.pro/news/wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented.md", "text": "https://wpnews.pro/news/wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented.txt", "jsonld": "https://wpnews.pro/news/wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented.jsonld"}}