cd /news/large-language-models/wiring-ios-coreml-to-a-quantized-on-… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-141585] src=dev.to β†— pub= topic=large-language-models verified=true sentiment=↑ positive

Wiring iOS CoreML to a Quantized On-Device Reranker for Retrieval-Augmented Generation

A developer outlined a two-stage on-device retrieval-augmented generation pipeline for iPhone that pairs an INT8-quantized MiniLM-L6 bi-encoder for approximate retrieval with a CoreML cross-encoder (MiniLM-L12) that rescores the top candidates, targeting an 80–120ms end-to-end latency budget on A16/A17 chips. The writeup reports that INT8 quantization costs under 0.5% NDCG on the bi-encoder and about 1.2% on the cross-encoder, while INT4 drops roughly 4.8% NDCG, and that reusing query-side KV states across candidates is a larger latency win than quantization alone.

by read5 min views4 publishedSep 29, 2026
---
title: "CoreML Cross-Encoder Reranking: On-Device RAG on iPhone"
published: true
description: "Build a two-stage on-device RAG pipeline on iPhone using CoreML β€” INT8-quantized bi-encoder retrieval, cross-encoder reranking, and A16/A17 latency budgets explained."
tags: ios, swift, mobile, architecture
canonical_url: https://mvpfactory.co/blog/coreml-cross-encoder-reranking-on-device-rag-iphone
---

## What We Are Building

By the end of this walkthrough, you will have a production-grade two-stage RAG pipeline running entirely on-device. Stage one: fast approximate retrieval with a quantized bi-encoder. Stage two: a CoreML cross-encoder that rescores the top-k candidates. The latency budget on A16/A17 chips is roughly 80–120ms end-to-end before users notice lag. Let me show you the pattern I use to get there.

## Prerequisites

- Xcode 15+
- `coremltools` 7.x installed in your Python environment
- A pre-trained bi-encoder (MiniLM-L6) and cross-encoder (MiniLM-L12) in PyTorch
- Basic familiarity with CoreML model conversion

---

## The Two-Stage Architecture

Here is the mental model worth internalising before writing a single line of code:

Query β†’ [Bi-Encoder] β†’ embedding β†’ ANN search (top-k=50)

                                      ↓

                          [Cross-Encoder] β†’ rescore top-k=10

                                      ↓

                              Final ranked results
The bi-encoder runs **once per query**. The cross-encoder runs **k times**. This asymmetry is the entire reason the two-stage model exists β€” cross-encoders are dramatically more accurate but cannot scale to full corpus retrieval.

---

## Step 1: Quantization β€” Get the Numbers Right

Most teams get this wrong: they quantize both stages with the same strategy and wonder why quality collapses. The bi-encoder and cross-encoder have different sensitivity profiles.

| Model Stage | FP32 Size | INT8 Size | Latency (A17) | Quality Drop |
|---|---|---|---|---|
| Bi-encoder (MiniLM-L6) | 90 MB | 24 MB | 8 ms | < 0.5% NDCG |
| Cross-encoder (MiniLM-L12) | 180 MB | 47 MB | 34 ms/candidate | ~1.2% NDCG |
| Cross-encoder (INT4 aggressive) | 180 MB | 23 MB | 19 ms/candidate | ~4.8% NDCG |

INT8 on the bi-encoder is nearly free. On the cross-encoder, INT8 is acceptable; INT4 starts hurting meaningful recall. Stay at INT8 for cross-encoders in production.

Export with CoreML Tools like this:

python

import coremltools as ct

mlmodel = ct.convert(

traced_model,

compute_precision=ct.precision.FLOAT16,

compute_units=ct.ComputeUnit.ALL  # uses ANE + GPU

)

mlmodel.save("CrossEncoder.mlpackage")

Target `ct.ComputeUnit.ALL` to push attention layers onto the Neural Engine. On A16 and A17, this cuts cross-encoder latency by roughly 40% versus CPU-only execution.

---

## Step 2: KV-Cache Reuse Across Candidates

Here is the gotcha that will save you hours. The biggest latency win in cross-encoder scoring is not quantization β€” it is KV-cache reuse. For a given query, the query-side key/value representations are identical across all k candidates. You compute them once.

swift

// Cache query KV states once per query

let queryKV = crossEncoder.encodeQuery(queryTokens)

let scores = candidates.map { doc in

crossEncoder.scoreWithCachedQuery(queryKV, docTokens: doc.tokens)

}

CoreML does not expose KV-cache injection directly for transformer models. You have two options: split the model at the cross-attention boundary and pass cached states manually, or use ONNX Runtime with a custom CoreML execution provider. Go with the split-model approach. It avoids the ONNX Runtime dependency, stays entirely within the Apple toolchain, and is straightforward to debug with Instruments.

---

## Step 3: Respect the Memory Pressure Ceiling

The constraint nobody mentions until production: the Neural Engine on A16/A17 shares memory bandwidth with the GPU, CPU, and camera ISP. Under sustained load, the kernel will thermally throttle ANE frequency within 60–90 seconds.

The practical ceiling:

- **k = 10–20 candidates**: Safe. Stays below 200ms total, no throttle.
- **k = 30–50 candidates**: Borderline. Latency climbs to 400–600ms. Sustained sessions trigger throttle at ~45s.
- **k > 50**: Avoid entirely on-device unless you batch overnight with background processing.

For interactive sessions, monitor thermal state and shed load before the system does it for you:

swift

let thermal = NSProcessInfo.processInfo.thermalState

if thermal == .serious || thermal == .critical {

// reduce k or defer reranking

}

For sustained background workloads, schedule reranking via `BGProcessingTaskRequest`, which the OS grants during charging when thermals are favorable.

---

## Gotchas

**INT4 looks tempting β€” resist it.** The 2Γ— size reduction is not worth the ~5% NDCG regression on cross-encoders. Validate against your own eval set (200–500 query/document pairs, scored with `pytrec_eval`) before committing to any quantization strategy.

**The working set math matters.** For a 12-layer cross-encoder scoring 20 candidates of 512 tokens, you are looking at ~310MB active β€” fine. At 50 candidates, you are pushing 750MB and competing with the OS.

**Verify quality regressions before shipping.** A 5-minute offline eval loop with `pytrec_eval` catches regressions between FP32 and quantized outputs. The docs do not mention this, but skipping it is how quality silently degrades in production.

Speaking of sustained focus work β€” if you are doing long CoreML optimization sessions, [HealthyDesk](https://play.google.com/store/apps/details?id=com.healthydesk) is worth keeping open in the background. Thermal throttling your chips and your own posture in the same session is a bad combo.

---

## Conclusion

Three things worth remembering:

1. **Use INT8 for both stages, not INT4.** Validate against your eval set before committing.
2. **Implement query-side KV-cache reuse via the split-model approach.** This cuts reranking latency by 30–45% on A16/A17 with no quality cost.
3. **Cap candidate count at k=20 for interactive sessions.** Schedule anything heavier via `BGProcessingTaskRequest`.

The two-stage pipeline is the right architecture for on-device RAG. Get the quantization strategy and the memory ceiling right, and you will stay comfortably inside the 80–120ms budget users never notice.
── more in #large-language-models 4 stories Β· sorted by recency
── more on @coreml 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/wiring-ios-coreml-to…] indexed:0 read:5min 2026-09-29 Β· β€”