# Wiring iOS CoreML to a Quantized On-Device Reranker for Retrieval-Augmented Generation

> Source: <https://dev.to/software_mvp-factory/wiring-ios-coreml-to-a-quantized-on-device-reranker-for-retrieval-augmented-generation-1o08>
> Published: 2026-09-29 08:38:42+00:00



```
---
title: "CoreML Cross-Encoder Reranking: On-Device RAG on iPhone"
published: true
description: "Build a two-stage on-device RAG pipeline on iPhone using CoreML — INT8-quantized bi-encoder retrieval, cross-encoder reranking, and A16/A17 latency budgets explained."
tags: ios, swift, mobile, architecture
canonical_url: https://mvpfactory.co/blog/coreml-cross-encoder-reranking-on-device-rag-iphone
---

## What We Are Building

By the end of this walkthrough, you will have a production-grade two-stage RAG pipeline running entirely on-device. Stage one: fast approximate retrieval with a quantized bi-encoder. Stage two: a CoreML cross-encoder that rescores the top-k candidates. The latency budget on A16/A17 chips is roughly 80–120ms end-to-end before users notice lag. Let me show you the pattern I use to get there.

## Prerequisites

- Xcode 15+
- `coremltools` 7.x installed in your Python environment
- A pre-trained bi-encoder (MiniLM-L6) and cross-encoder (MiniLM-L12) in PyTorch
- Basic familiarity with CoreML model conversion

---

## The Two-Stage Architecture

Here is the mental model worth internalising before writing a single line of code:
```

Query → [Bi-Encoder] → embedding → ANN search (top-k=50)

                                          ↓

                              [Cross-Encoder] → rescore top-k=10

                                          ↓

                                  Final ranked results

```
The bi-encoder runs **once per query**. The cross-encoder runs **k times**. This asymmetry is the entire reason the two-stage model exists — cross-encoders are dramatically more accurate but cannot scale to full corpus retrieval.

---

## Step 1: Quantization — Get the Numbers Right

Most teams get this wrong: they quantize both stages with the same strategy and wonder why quality collapses. The bi-encoder and cross-encoder have different sensitivity profiles.

| Model Stage | FP32 Size | INT8 Size | Latency (A17) | Quality Drop |
|---|---|---|---|---|
| Bi-encoder (MiniLM-L6) | 90 MB | 24 MB | 8 ms | < 0.5% NDCG |
| Cross-encoder (MiniLM-L12) | 180 MB | 47 MB | 34 ms/candidate | ~1.2% NDCG |
| Cross-encoder (INT4 aggressive) | 180 MB | 23 MB | 19 ms/candidate | ~4.8% NDCG |

INT8 on the bi-encoder is nearly free. On the cross-encoder, INT8 is acceptable; INT4 starts hurting meaningful recall. Stay at INT8 for cross-encoders in production.

Export with CoreML Tools like this:
```

python

import coremltools as ct

mlmodel = ct.convert(

    traced_model,

    compute_precision=ct.precision.FLOAT16,

    compute_units=ct.ComputeUnit.ALL  # uses ANE + GPU

)

mlmodel.save("CrossEncoder.mlpackage")

```
Target `ct.ComputeUnit.ALL` to push attention layers onto the Neural Engine. On A16 and A17, this cuts cross-encoder latency by roughly 40% versus CPU-only execution.

---

## Step 2: KV-Cache Reuse Across Candidates

Here is the gotcha that will save you hours. The biggest latency win in cross-encoder scoring is not quantization — it is KV-cache reuse. For a given query, the query-side key/value representations are identical across all k candidates. You compute them once.
```

swift

// Cache query KV states once per query

let queryKV = crossEncoder.encodeQuery(queryTokens)

let scores = candidates.map { doc in

    crossEncoder.scoreWithCachedQuery(queryKV, docTokens: doc.tokens)

}

```
CoreML does not expose KV-cache injection directly for transformer models. You have two options: split the model at the cross-attention boundary and pass cached states manually, or use ONNX Runtime with a custom CoreML execution provider. Go with the split-model approach. It avoids the ONNX Runtime dependency, stays entirely within the Apple toolchain, and is straightforward to debug with Instruments.

---

## Step 3: Respect the Memory Pressure Ceiling

The constraint nobody mentions until production: the Neural Engine on A16/A17 shares memory bandwidth with the GPU, CPU, and camera ISP. Under sustained load, the kernel will thermally throttle ANE frequency within 60–90 seconds.

The practical ceiling:

- **k = 10–20 candidates**: Safe. Stays below 200ms total, no throttle.
- **k = 30–50 candidates**: Borderline. Latency climbs to 400–600ms. Sustained sessions trigger throttle at ~45s.
- **k > 50**: Avoid entirely on-device unless you batch overnight with background processing.

For interactive sessions, monitor thermal state and shed load before the system does it for you:
```

swift

let thermal = NSProcessInfo.processInfo.thermalState

if thermal == .serious || thermal == .critical {

    // reduce k or defer reranking

}

```
For sustained background workloads, schedule reranking via `BGProcessingTaskRequest`, which the OS grants during charging when thermals are favorable.

---

## Gotchas

**INT4 looks tempting — resist it.** The 2× size reduction is not worth the ~5% NDCG regression on cross-encoders. Validate against your own eval set (200–500 query/document pairs, scored with `pytrec_eval`) before committing to any quantization strategy.

**The working set math matters.** For a 12-layer cross-encoder scoring 20 candidates of 512 tokens, you are looking at ~310MB active — fine. At 50 candidates, you are pushing 750MB and competing with the OS.

**Verify quality regressions before shipping.** A 5-minute offline eval loop with `pytrec_eval` catches regressions between FP32 and quantized outputs. The docs do not mention this, but skipping it is how quality silently degrades in production.

Speaking of sustained focus work — if you are doing long CoreML optimization sessions, [HealthyDesk](https://play.google.com/store/apps/details?id=com.healthydesk) is worth keeping open in the background. Thermal throttling your chips and your own posture in the same session is a bad combo.

---

## Conclusion

Three things worth remembering:

1. **Use INT8 for both stages, not INT4.** Validate against your eval set before committing.
2. **Implement query-side KV-cache reuse via the split-model approach.** This cuts reranking latency by 30–45% on A16/A17 with no quality cost.
3. **Cap candidate count at k=20 for interactive sessions.** Schedule anything heavier via `BGProcessingTaskRequest`.

The two-stage pipeline is the right architecture for on-device RAG. Get the quantization strategy and the memory ceiling right, and you will stay comfortably inside the 80–120ms budget users never notice.
```


