# Replacing an LLM With a Lookup

> Source: <https://dev.to/ataakkyn/replacing-an-llm-with-a-lookup-198l>
> Published: 2026-09-30 08:08:23+00:00

On a B2B outbound product, one pipeline step used an LLM with web search to find a company's website and phone number from its name. Swapping it for a deterministic Places API lookup took website completeness from 34% to 81%, phone to 91%, and that step's LLM cost to zero. This is the reasoning behind the swap, and how to spot the same pattern in your own pipeline.

The problem

The product was an outbound sales system for freight forwarders. Given a company name and a country, it had to find that company's website and phone number before anything else could run. No website meant no enrichment, no outreach, no lead.

We started with an LLM that had web search, and it looked like the obvious fit. The input is messy: trading names, legal suffixes, transliterations, the same company spelled three ways across three sources. The web is unstructured. A model that can search and read pages should absorb exactly that kind of fuzziness. It also took about twenty minutes to wire up.

For the first hundred leads it looked like it worked.

Where it broke down

Four failures, roughly in the order they hurt.

It invented domains. Not often, but often enough. A plausible URL for a company that did not own it. Because a fabricated answer is shaped exactly like a correct one, nothing downstream caught it. We found out when outreach started reaching the wrong companies.

The same input gave different answers. Re-run a lead and the domain could change. That makes the step untestable: there is no fixture to assert against, and no way to separate a regression from ordinary variance.

Nothing was cacheable. The same company across three campaigns meant three paid lookups.

Cost scaled with lead volume, which is the one number you want to grow. Every improvement at the top of the funnel made this step linearly more expensive, forever.

The completeness figure was the quiet one. Only 34% of leads came out with a usable website, and we had filed that under data problem — small freight forwarders, thin web presence, nothing to be done. It was not a data problem.

The question that decides it

Here is the test I now run against every model call in a pipeline:

Does this step need judgement, or does it need an index?

An LLM is a reasoning engine that happens to have memorised a lot. When the task is genuinely open — summarise this thread, decide whether this reply is a rejection, draft a sentence in this register — you want the reasoning, and non-determinism is a fair price for it.

Finding a company's website is not that. There is one correct answer, it is a fact about the world, and somebody already maintains an index of it. Asking a model to search for it is asking it to reconstruct, unreliably and at cost, a lookup that a database performs exactly for a fraction of a cent.

The giveaway is the acceptance test. If expected == actual is a sensible assertion for the step, it is a lookup, and a model is the wrong tool for it.

The swap

The replacement was the Google Places API, but the API call was the easy part. What made it work was the layer around it, and that layer is what usually gets skipped.

Normalisation. Raw company names carry noise the index does not: legal suffixes (Ltd, A.Ş., GmbH, SARL), punctuation, casing, transliteration. Names are normalised before the query, not after the results come back.

Query construction. The name alone is too weak a key. Name plus city plus country narrows the search enough that the top candidate is usually right. Without the geography, the index happily returns a real business with a similar name in the wrong country — which is the same failure we had just removed.

Match scoring. The API returns candidates, not an answer. Each candidate is scored on token overlap with the normalised name, geographic distance, and business-type agreement. Below a threshold, the step returns nothing.

That last rule matters more than the accuracy figure. Returning nothing is a state the rest of the system can handle: the lead is flagged and routed to manual review. Returning a confident wrong answer is not a state anything can handle, because nothing downstream knows to doubt it.

Caching. Results are keyed on the normalised name plus country and stored. The second campaign against the same company costs nothing.

Results

Measured across the same lead set, before and after the swap:

```
LLM + web search    Places lookup
```

Website completeness    34% 81%

LLM cost for this step  per lead, unbounded zero

Same input, same output no  yes

Cacheable   no  yes

Testable against fixtures   no  yes

Phone completeness reached 91% after the swap. I have no honest before-figure to compare it against: the LLM path frequently returned no phone number at all, so the metric was never tracked properly until there was something worth tracking.

The number I did not expect was the last row. Removing the model from this step did not just make it cheaper — it made it a thing I could write tests for, which changed how quickly everything downstream of it could be changed.

What stayed on the model

This is not an argument against LLMs in pipelines. Two steps in the same system kept theirs, because both fail the index test.

Drafting outreach copy. There is no correct answer to look up. The model writes, a human approves, and nothing is sent without that approval.

Classifying inbound replies. Deciding whether a reply is interest, a rejection, an out-of-office or a request to be removed is judgement over free text. An index cannot do it.

Both are open-ended tasks where non-determinism is acceptable and a human sits in the loop. That is the shape of work worth paying a model for.

The budget guard

The remaining calls got a monthly spend cap. When the cap is reached, generation stops and the rest of the system keeps running — leads still enrich, replies still classify against the cached model outputs, the panel still works. Only new generation pauses, and it pauses loudly.

The design rule is that a cost limit should degrade one feature, not take the product down. A guard that halts everything is an outage you built for yourself, and the second time it fires someone will remove it rather than fix it.

Audit your own pipeline

Go through every model call in your system and ask five questions. Any step that answers this way is a lookup wearing a model's clothes:

Could you write expected == actual as its test?

Does a public API, a database or a file already hold the answer?

Would two runs on the same input disagree, and would that be a bug rather than a variation?

Is its cost proportional to your input volume rather than to your user count?

When it is wrong, does the output still look right?

The last one is the dangerous question. A step that fails loudly gets fixed in a week. A step that fails silently and plausibly stays in production for months, and quietly poisons everything downstream of it.

Reproducing this

The repo is at github.com/ataakkoyun68-png/llm-vs-lookup. It runs both strategies over the same synthetic dataset and prints the comparison table. Node 22, no install, no API keys:

.

├─ data/

│  ├─ companies.csv          # 24 synthetic leads, messy names, 6 countries

│  ├─ ground-truth.csv       # what is actually true; only the benchmark sees it

│  └─ places-index.json      # the index, with name variants and decoys

├─ src/

│  ├─ types.ts               # the interface both strategies sit behind

│  ├─ normalise.ts           # suffix stripping, transliteration, cache keys

│  ├─ score.ts               # token overlap, geography, type, abstain threshold

│  ├─ budget-guard.ts        # monthly cap; degrade one feature, not all

│  ├─ strategies/

│  │  ├─ llm.ts              # simulated failure modes, explicit rates

│  │  └─ places.ts           # deterministic lookup + match scoring + cache

│  └─ benchmark.ts           # runs both twice, measures, prints the table

└─ README.md

Both strategies sit behind one interface, so the benchmark is a fair comparison rather than a rhetorical one. Swap in your own dataset and the numbers will differ from mine — the point is the method, not my percentages.

One thing to be clear about: the LLM strategy in that repo is a simulation. It does not call a model. It reproduces the four failure modes above with the rates written down as explicit parameters so you can argue with them, and the dataset is synthetic. The production numbers quoted at the top came from the real system; the repo demonstrates the architecture.
