What Makes LLM Testing Different? A developer argues that LLM testing requires a different approach than traditional automated testing, because generated responses can be semantically correct while differing in wording. The developer recommends combining deterministic assertions—such as checking for the presence of "5" or "five"—with semantic evaluations like relevance and groundedness scores, and notes that failures may stem from either outdated retrieval or model hallucination. Traditional automated testing is usually straightforward: provide an input, define the expected output, and compare the two. If I am testing a function that calculates a discount, the same input should consistently produce the same result. That makes assertions simple. LLM applications behave differently. Imagine we’re testing an employee assistant against a policy that says employees can carry over a maximum of 5 vacation days. We ask: How many unused vacation days can I carry over? The model could answer “You can carry over up to five unused vacation days” or “The maximum vacation carryover is five days.” Both are correct, even though the strings are completely different. This makes an assertion like the following too restrictive: expected = "You can carry over up to five unused vacation days." response = ask llm "How many unused vacation days can I carry over?" assert response == expected The test may fail simply because the model chose different words. Instead, we can start by validating the things that are deterministic and then evaluate the semantic quality separately. python def test vacation policy : question = "How many unused vacation days can I carry over?" context = """ Employees may carry over a maximum of 5 unused vacation days into the next year. """ response = ask llm question, context Deterministic checks assert response is not None assert len response.strip 0 assert "5" in response or "five" in response.lower Semantic checks assert evaluate relevance question, response = 0.8 assert evaluate groundedness context, response = 0.8 This is where LLM testing starts to look different. The first few assertions behave like normal automated tests. But relevance asks whether the response actually answers the user’s question, while groundedness asks whether the answer is supported by the supplied policy. Suppose our test suddenly receives “Employees can carry over 10 vacation days.” The API might still return 200, the response might be perfectly formatted, and latency might be excellent—but the application is wrong. We now need to determine whether retrieval supplied an outdated policy or whether the correct five-day policy was retrieved and the LLM hallucinated the number ten. That’s why I don’t think LLM testing replaces traditional testing. It adds another layer to it. We still test APIs, schemas, permissions, latency, tool calls, and integrations deterministically. But for generated responses, we also need to evaluate qualities such as relevance, groundedness, hallucination, and consistency. A useful way to think about the difference is: Traditional testing asks: “Did I get the expected output?” LLM testing asks: “Is this output correct and trustworthy, even if it isn’t exactly what I expected?” That small change in the definition of “correct” is what makes testing LLM applications fundamentally different.