# Cache Me If You Can: 10x Faster Start on My Laptop, a Bigger Bill on Claude

> Source: <https://dev.to/shivangb237/cache-me-if-you-can-10x-faster-start-on-my-laptop-a-bigger-bill-on-claude-g5o>
> Published: 2026-10-08 20:10:26+00:00

**TL;DR** Same RAG pipeline, two setups: a MacBook Air running a local 4B model, and Claude Sonnet 5.5. Four kinds of caching, switched on and off. Every number is from logged runs.

- ⚡
**Laptop:** with the retrieval cache and prompt caching working together on questions that share paragraphs, the model started answering about 10x sooner. Time to first token (how long before it begins answering) fell from 3–4 s to 0.3 s.- 💸
**Claude:** prompt caching raised the bill 18–25%. Making prompts repeat on purpose fixed it, but only for questions clustered around the same paragraphs (23% cheaper).- ⚠️
**Similarity caches** can serve confidently wrong answers. I have receipts.

I measured the four caches people put in front of an LLM: exact-match, semantic, retrieval and prompt (prefix) caching. Same RAG pipeline (the model answers using passages looked up first), scored against SQuAD, a public question-answering dataset. The surprise: **the same cache solved a different problem on each machine.**

**Who this is for:** anyone adding a semantic or prompt cache to a RAG pipeline because a tutorial said to.

*Where a question can be answered before it reaches the model.*

| Cache | 🖥️ Laptop (Qwen3 4B) | ☁️ Claude Sonnet 5.5 | 😬 The catch | 
|---|---|---|---|
| **Exact-match** | A hit: 0.1–0.2 ms instead of 4.6–5.7 s | 15% fewer calls, 15% lower cost at 15% repeats | Served the old number for 29 of 30 questions after the data changed (0 of 30 with invalidation by data version) | 
| **Semantic** (0.90) | 34% answered from cache, median time −10% | 39% answered from cache, 39% cheaper | 4 wrong answers per run on Claude (0 of 24 near-misses on the laptop) | 
| **Retrieval** | Prompt processed in under 1 s: 19% → 55% of requests | Makes prompt caching pay (Round 2) | Answer quality (F1) −0.014 locally, −0.02 to −0.03 on Claude | 
| **Prompt (prefix)** | First token 3–4 s → 0.3 s (with retrieval cache) | +18% cost alone, −23% with retrieval cache (clustered) | No speedup on Claude; +20–25% on diverse questions | 

With no caches, a median question took 4.7 seconds, and the first token (the first piece of the answer) alone took 4.4. Prompts averaged 942 tokens and answers 8.8: the model spends its time reading, not writing.

Ollama, which runs the model, already caches recent prompts in up to 8 GiB of RAM, about 60 prompts of this size. A repeated 885-token prompt took 29 ms instead of 2,770 ms.

The catch: only the start of a prompt is reusable, up to the first difference. A RAG prompt starts with a fixed instruction (about 6% of it), then passages that change per question. Across 200 questions touching about 190 paragraphs, only about 10% of prompt tokens were reused.

A bigger cache was not the fix. Making prompts repeat was. A retrieval cache remembers which passages a question found and hands the same passages, in the same order, to similar questions, so the passage block matches what the model has already read. Order matters: the same passages in a different order break the match.

On 100 shuffled questions about 20 paragraphs, with a similarity threshold of 0.65 (how alike two questions must be, on a 0 to 1 scale, to share passages), requests whose prompt was processed in under a second rose from 19% to 55%. The median time to first token fell from 3-4 seconds to about 0.3.

*Only the loosest threshold, 0.65, made a clear difference.*

Claude's prompt caching is explicit: you mark a breakpoint after the passages, which tells the API that everything before it can be reused. Writing costs 25% more than normal input; reading costs about 90% less. So it only pays if enough gets read back: cost is roughly `1 + 0.25 x (share written) - 0.9 x (share read)` times the normal prompt cost, with shares counted in prompt tokens.

I ran two workloads of 100 questions: clustered (the same 20 paragraphs) and diverse. Cost per 100 questions, clustered:

*Caching on Claude only paid when prompts repeated on purpose.*

On the diverse workload nothing repeated: every cached setup cost 20-25% more than no caching.

Speed did not improve either. Writing a cache entry made the median time to first token 20-28% slower, and reads did not speed it up. Claude already answers in about 1.4 seconds. Locally the cache buys time. On the API it buys money, and only if the start of your prompts repeats.

A semantic cache answers a question with the saved answer of a *similar* one, and never asks the model. Fast, free, and sometimes wrong. These are real near-misses (same paragraph, different question) from my Claude run, served from the semantic cache for $0:

| Asked | Cache answered | Gold answer | Time | 
|---|---|---|---|
| Who is the vice-chair of the IPCC? | Hoesung Lee | Ismail El Gizouli | 8 ms | 
| Who was the first chair of the IPCC? | Hoesung Lee | Bert Bolin | 11 ms | 
| Red ribbons in the logo were used to represent which division of ABC? | ABC News | the entertainment division | 13 ms | 

Every run of 201 requests had 4 of these, all on near-misses, while the cache cut cost by 39%. The worst offenders scored 0.93–0.98 similarity: a swapped word (“who” vs “when”, “before” vs “after”) barely moves an embedding. **Similarity is not equivalence.**

*Replay of a real run: red marks a wrong answer served from cache.*

Replay the whole run, request by request, in the **[live demo](https://shivb237.github.io/cache-rag-gateway/)**.

☁️ Bonus: I shipped it to AWS (click to expand)

To check the gateway outside my laptop, I ran it as a container on ECS Fargate (AWS's service for running containers without managing servers) in Sydney, building every piece by hand.

The image bakes in the search index and embedding model and runs as a non-root user. My first build was 2.79 GB because `chown -R` copied every installed file into a new layer. `USER` before the install plus `COPY --chown` brought it to 1.62 GB (498 MB compressed).

The setup is deliberately small:

`x-api-key` header.
What I would have missed: registries and login tokens are per region (the wrong one gives "no basic auth credentials"); my push identity could only touch ECR, so creating a repository or log group was refused, which is least privilege working; and ECS ignores the Dockerfile's `HEALTHCHECK`, so it must go in the task definition.

One test on the running task: the first request took 2,359 ms and cost $0.0035. The same question again came back from the exact cache in 0 ms for $0. No key got a 401 (unauthorized). Times are measured inside the container. This is a test deployment, not production: one task, no HTTPS, and a replaced task starts with empty caches and a reset spend counter. I stopped the task afterwards so it would not keep billing.

``` bash
$ curl -s -X POST http://:8080/ask ... -d '{"question":"Who was the NFL Commissioner in early 2012?"}'

{"answer":"Roger Goodell","refused":false,"cacheLayer":null,"latencyMs":2359,"costUsd":0.0035425}

$ (same request again)

{"answer":"Roger Goodell","refused":false,"cacheLayer":"exact","latencyMs":0,"costUsd":0}
```

Claude's newest models reject the temperature setting (the control that makes output repeatable), so runs vary. To measure how much, every experiment includes an identical copy of the control (no caches). Two copies of one setup disagreed on 25% of answers and differed by 0.013 in F1 (a 0-1 score for how closely an answer matches the correct one), so I treat smaller differences as noise.

My setup was wrong more than once. One control was accidentally helped by leftover entries in the model's prompt cache from an earlier test, so I discarded its first 10 requests. Claude refused 22 of 4,221 requests, all on one harmless biology question (and its paraphrase) about how a virus evades detection; I count refusals as wrong answers and never cache them.

The retrieval cache has a price: a question borrows passages found for a different one, lowering F1 by 0.014 locally and 0.02-0.03 on Claude. And my clustered workload covered only 20 paragraphs, which flatters reuse. One dataset, two models.

**What does each layer of LLM caching actually save, and what does it cost you in correctness?**

A small RAG gateway where every cache layer can be switched on or off, plus an evaluation harness that measures cost, latency, answer quality, wrong answers served and staleness against SQuAD gold answers. All results come from real runs logged in [`runs/`](https://github.com/Shivb237/cache-rag-gateway/runs/), on a local 4B model (Ollama) and on Claude Sonnet 5.5.

``` php
flowchart LR
  Q[question] --> E{exact cache}
  E -- hit --> A[answer]
  E -- miss --> S{semantic cache}
  S -- hit --> A
  S -- miss --> R[embed + retrieve top-k<br/>optional retrieval cache]
  R --> P[prompt: instruction, passages, question<br/>optional provider prompt cache]
  P --> M[(LLM)]
  M --> A
```

**Demo:** [replay 201 real requests](https://shivb237.github.io/cache-rag-gateway/) and see which layer answered, the latency and the cost, for each cache setup. The page replays logged values from [`runs/ablation-claude.jsonl`](https://github.com/Shivb237/cache-rag-gateway/runs/ablation-claude.jsonl) (`npm run replay`…

Every log is in the repo's `runs/` folder, and the demo replays them: [https://shivb237.github.io/cache-rag-gateway/](https://shivb237.github.io/cache-rag-gateway/)

💬 **Your turn:** has prompt caching or a semantic cache saved you money, time, or neither on a real endpoint? Tell me in the comments, especially if it contradicts mine.

*Data: SQuAD v1.1 (CC BY-SA 4.0).*
