{"slug": "replacing-an-llm-with-a-lookup", "title": "Replacing an LLM With a Lookup", "summary": "A developer building a B2B outbound sales pipeline for freight forwarders replaced an LLM-with-web-search step for finding company websites and phone numbers with a deterministic Google Places API lookup, raising website completeness from 34% to 81% and phone coverage to 91% while cutting that step's LLM cost to zero. The writeup details four failure modes of the LLM approach — fabricated domains, non-deterministic answers, no cacheability, and cost scaling with lead volume — and describes the normalization, geography-aware query construction, and match-scoring layer that made the lookup reliable.", "body_md": "On a B2B outbound product, one pipeline step used an LLM with web search to find a company's website and phone number from its name. Swapping it for a deterministic Places API lookup took website completeness from 34% to 81%, phone to 91%, and that step's LLM cost to zero. This is the reasoning behind the swap, and how to spot the same pattern in your own pipeline.\n\nThe problem\n\nThe product was an outbound sales system for freight forwarders. Given a company name and a country, it had to find that company's website and phone number before anything else could run. No website meant no enrichment, no outreach, no lead.\n\nWe started with an LLM that had web search, and it looked like the obvious fit. The input is messy: trading names, legal suffixes, transliterations, the same company spelled three ways across three sources. The web is unstructured. A model that can search and read pages should absorb exactly that kind of fuzziness. It also took about twenty minutes to wire up.\n\nFor the first hundred leads it looked like it worked.\n\nWhere it broke down\n\nFour failures, roughly in the order they hurt.\n\nIt invented domains. Not often, but often enough. A plausible URL for a company that did not own it. Because a fabricated answer is shaped exactly like a correct one, nothing downstream caught it. We found out when outreach started reaching the wrong companies.\n\nThe same input gave different answers. Re-run a lead and the domain could change. That makes the step untestable: there is no fixture to assert against, and no way to separate a regression from ordinary variance.\n\nNothing was cacheable. The same company across three campaigns meant three paid lookups.\n\nCost scaled with lead volume, which is the one number you want to grow. Every improvement at the top of the funnel made this step linearly more expensive, forever.\n\nThe completeness figure was the quiet one. Only 34% of leads came out with a usable website, and we had filed that under data problem — small freight forwarders, thin web presence, nothing to be done. It was not a data problem.\n\nThe question that decides it\n\nHere is the test I now run against every model call in a pipeline:\n\nDoes this step need judgement, or does it need an index?\n\nAn LLM is a reasoning engine that happens to have memorised a lot. When the task is genuinely open — summarise this thread, decide whether this reply is a rejection, draft a sentence in this register — you want the reasoning, and non-determinism is a fair price for it.\n\nFinding a company's website is not that. There is one correct answer, it is a fact about the world, and somebody already maintains an index of it. Asking a model to search for it is asking it to reconstruct, unreliably and at cost, a lookup that a database performs exactly for a fraction of a cent.\n\nThe giveaway is the acceptance test. If expected == actual is a sensible assertion for the step, it is a lookup, and a model is the wrong tool for it.\n\nThe swap\n\nThe replacement was the Google Places API, but the API call was the easy part. What made it work was the layer around it, and that layer is what usually gets skipped.\n\nNormalisation. Raw company names carry noise the index does not: legal suffixes (Ltd, A.Ş., GmbH, SARL), punctuation, casing, transliteration. Names are normalised before the query, not after the results come back.\n\nQuery construction. The name alone is too weak a key. Name plus city plus country narrows the search enough that the top candidate is usually right. Without the geography, the index happily returns a real business with a similar name in the wrong country — which is the same failure we had just removed.\n\nMatch scoring. The API returns candidates, not an answer. Each candidate is scored on token overlap with the normalised name, geographic distance, and business-type agreement. Below a threshold, the step returns nothing.\n\nThat last rule matters more than the accuracy figure. Returning nothing is a state the rest of the system can handle: the lead is flagged and routed to manual review. Returning a confident wrong answer is not a state anything can handle, because nothing downstream knows to doubt it.\n\nCaching. Results are keyed on the normalised name plus country and stored. The second campaign against the same company costs nothing.\n\nResults\n\nMeasured across the same lead set, before and after the swap:\n\n```\nLLM + web search    Places lookup\n```\n\nWebsite completeness    34% 81%\n\nLLM cost for this step  per lead, unbounded zero\n\nSame input, same output no  yes\n\nCacheable   no  yes\n\nTestable against fixtures   no  yes\n\nPhone completeness reached 91% after the swap. I have no honest before-figure to compare it against: the LLM path frequently returned no phone number at all, so the metric was never tracked properly until there was something worth tracking.\n\nThe number I did not expect was the last row. Removing the model from this step did not just make it cheaper — it made it a thing I could write tests for, which changed how quickly everything downstream of it could be changed.\n\nWhat stayed on the model\n\nThis is not an argument against LLMs in pipelines. Two steps in the same system kept theirs, because both fail the index test.\n\nDrafting outreach copy. There is no correct answer to look up. The model writes, a human approves, and nothing is sent without that approval.\n\nClassifying inbound replies. Deciding whether a reply is interest, a rejection, an out-of-office or a request to be removed is judgement over free text. An index cannot do it.\n\nBoth are open-ended tasks where non-determinism is acceptable and a human sits in the loop. That is the shape of work worth paying a model for.\n\nThe budget guard\n\nThe remaining calls got a monthly spend cap. When the cap is reached, generation stops and the rest of the system keeps running — leads still enrich, replies still classify against the cached model outputs, the panel still works. Only new generation pauses, and it pauses loudly.\n\nThe design rule is that a cost limit should degrade one feature, not take the product down. A guard that halts everything is an outage you built for yourself, and the second time it fires someone will remove it rather than fix it.\n\nAudit your own pipeline\n\nGo through every model call in your system and ask five questions. Any step that answers this way is a lookup wearing a model's clothes:\n\nCould you write expected == actual as its test?\n\nDoes a public API, a database or a file already hold the answer?\n\nWould two runs on the same input disagree, and would that be a bug rather than a variation?\n\nIs its cost proportional to your input volume rather than to your user count?\n\nWhen it is wrong, does the output still look right?\n\nThe last one is the dangerous question. A step that fails loudly gets fixed in a week. A step that fails silently and plausibly stays in production for months, and quietly poisons everything downstream of it.\n\nReproducing this\n\nThe repo is at github.com/ataakkoyun68-png/llm-vs-lookup. It runs both strategies over the same synthetic dataset and prints the comparison table. Node 22, no install, no API keys:\n\n.\n\n├─ data/\n\n│  ├─ companies.csv          # 24 synthetic leads, messy names, 6 countries\n\n│  ├─ ground-truth.csv       # what is actually true; only the benchmark sees it\n\n│  └─ places-index.json      # the index, with name variants and decoys\n\n├─ src/\n\n│  ├─ types.ts               # the interface both strategies sit behind\n\n│  ├─ normalise.ts           # suffix stripping, transliteration, cache keys\n\n│  ├─ score.ts               # token overlap, geography, type, abstain threshold\n\n│  ├─ budget-guard.ts        # monthly cap; degrade one feature, not all\n\n│  ├─ strategies/\n\n│  │  ├─ llm.ts              # simulated failure modes, explicit rates\n\n│  │  └─ places.ts           # deterministic lookup + match scoring + cache\n\n│  └─ benchmark.ts           # runs both twice, measures, prints the table\n\n└─ README.md\n\nBoth strategies sit behind one interface, so the benchmark is a fair comparison rather than a rhetorical one. Swap in your own dataset and the numbers will differ from mine — the point is the method, not my percentages.\n\nOne thing to be clear about: the LLM strategy in that repo is a simulation. It does not call a model. It reproduces the four failure modes above with the rates written down as explicit parameters so you can argue with them, and the dataset is synthetic. The production numbers quoted at the top came from the real system; the repo demonstrates the architecture.", "url": "https://wpnews.pro/news/replacing-an-llm-with-a-lookup", "canonical_source": "https://dev.to/ataakkyn/replacing-an-llm-with-a-lookup-198l", "published_at": "2026-09-30 08:08:23+00:00", "updated_at": "2026-09-30 08:16:28.730844+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-agents", "developer-tools"], "entities": ["Google Places API"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/replacing-an-llm-with-a-lookup", "markdown": "https://wpnews.pro/news/replacing-an-llm-with-a-lookup.md", "text": "https://wpnews.pro/news/replacing-an-llm-with-a-lookup.txt", "jsonld": "https://wpnews.pro/news/replacing-an-llm-with-a-lookup.jsonld"}}