cd /news/ai-agents/where-does-the-llm-go-in-the-testing… · home › topics › ai-agents › article
[ARTICLE · art-147735] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Where does the LLM go in the Testing Pyramid?

A developer proposes extending the traditional testing pyramid to account for LLM-based features, arguing that LLM calls sit between service and UI layers because they are as slow as UI tests but asserted on like APIs. The approach splits the middle layer into tests that mock the LLM and tests that call it, with the latter using an LLM to orchestrate multi-turn conversations and an LLM-as-judge to evaluate transcripts, trading maintenance cost against execution cost.

by read5 min views2 publishedOct 8, 2026

We all know and love the Test Pyramid, but if your application uses AI, there's a problem: it's now possible to write Service layer tests that are slower and more expensive than the UI tests.

Archibald's Pizzeria has just forked over $X,000 to a slick enterprise vendor to bring them into the 21st century. In addition to ordering pizzas on the website and mobile app, customers can now call "Archie", a virtual chatbot receptionist that takes orders over the phone.

Now look at these three test cases, and tell me where they go in the pyramid:

Test 1 is a regular old E2E UI test against a webapp and belongs squarely at the top of the pyramid. I believe Test 2 belongs there as well - the phone connection is the "UI" of a voice agent and the only way to test it end-to-end is to call it and talk to it.

Test 3 is where things get tricky. It's an API test, and it avoids the UI being used in test 2, so it ought to go in the middle of the pyramid, right? However, in most cases it will be slower and more expensive than test 1, because it hits an LLM.

Let's alter the pyramid to reflect this. LLMs are as slow or slower than the UI but you call them and assert on their output like an API, so let's put them in between:

We now have two service layers: tests that use the LLM and tests that avoid it somehow. Sounds like we need a way to mock the LLM!

The good news is, the mock should be technically straightforward to implement, because the whole point of an inference model is to infer what the user wants, and we already know what an automated test wants. Consider the pizza example. If Archibald's only sells two kinds of pizza, the mock LLM is a simple mapping of user intent to tool calls:

function mockLLM(userInput, pizzaTool) {
  if (userInput.includes("pepperoni")) pizzaTool.orderPepperoni();
  if (userInput.includes("hawaiian")) pizzaTool.orderHawaiian();
}

The bad news is, it's pure technical debt. It will have to be updated every time the menu changes, let alone the agent's functional behavior. It will also be complex, assuming your real agent has inference decisions that lead to tree decisions that open more inference decisions. The mock LLM will end up looking like a traditional IVR conversational script tree - which is what the AI was supposed to replace!

So, if mocking the LLM gets you fast/free tests at the cost of tech debt, any tests that go through the LLM had better be tech debt free. How do we ensure that?

The answer, I believe, is very clear: use an LLM to orchestrate the test call and judge the outcome. Remember the definition of inference: if the model's purpose is to figure out how to do what the user wants, then "what the user wants" is the test case. Using an LLM to test your LLM allows you to write test cases like this:

callArchie("You are John Smith of 1234 Main St and
you want to order a pepperoni pizza. Verify that
the agent confirms your selection, tells you the
price, remembers to charge sales tax, and closes
the call by thanking you.")

This test is absolutely zero percent tech debt. You could change IVA vendors, change LLM models, rewrite the whole stack from Java to C# and then back to Java again, and this test would stand, as is.

What you need to build to get there is that callArchie function, which is non-trivial but certainly doable. It needs to orchestrate a multi-turn conversation between your agent and another model (or the same one with different prompts), and then pass the transcript to an LLM-as-judge to evaluate. The script you can probably one-shot, but plan to spend an afternoon getting the prompts dialed in.

We've split the middle layer of the pyramid into "mocks the LLM" and "calls the LLM", and we've discussed what you need to build once and what you need to maintain for each kind. I believe the easiest way to decide which tests go where is to consider each test as a trade-off between maintenance cost vs. execution cost.

How you decide which test goes where is up to you. To an old-school SDET, an API test suite that costs a dollar to run sounds insanely expensive, I know. But tech debt is costly too; if the mock LLM requires even an hour a month to keep it synced to the application's functional behavior, and you only run your tests a few times a week, the mock might not be worth it.

Before we conclude, it's worth revisiting the top of the pyramid to point out that if you end up building the LLM-as-judge service layer test, you can have a true end-to-end test by running the exact same test through a phone line. How you do that and how hard it is depends entirely on what your voice platform exposes:

From my experience testing voice agents and LLM-powered applications, deciding whether the LLM "counts" as part of the UI layer or the service layer depends entirely on how easily you can mock the LLM and how well you can test your service layer without going through it.

The Practical Test Pyramid - a deeper dive on the original model

Writing effective Voice Agent tests - how to phrase test goals for the inference layer

── more in #ai-agents 4 stories · sorted by recency
── more on @archibald's pizzeria 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/where-does-the-llm-g…] indexed:0 read:5min 2026-10-08 · —