# Presentation: Multi-Agent Patterns from Spotify’s AI Powered Advertising Platform

> Source: <https://www.infoq.com/presentations/spotify-multi-agent-ai-architecture/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global>
> Published: 2026-10-08 11:00:00+00:00

## Transcript

"Learn how to run AI at scale at QCon AI conference. Moving AI from experiments into production increases pressure on platforms and teams. Hear senior peers explain how to control inference latency, build safe agent guardrails, and standardize internal platforms. Plus an early bird price until May 12th. Register now".

**Pratik Rasam:** The ad that you just heard, the script, the audience, the targeting, all of that was built by an in-house production platform that we have within Spotify Ads Manager. If you're into advertising, you'll know that creating an ad involves copywriting, music generation. All of that is now not needed. All we need is just the advertiser to describe what they want in an ad. This is not a research talk, but these are actual production agents at scale running at Spotify Ads Manager. I will also be talking about the mistakes that we did while developing this platform.

## Background

My name is Pratik. I'm a senior engineer at Spotify Ads. Today I will be going over the platform that builds Ads AI. Ads AI is basically a multi-agentic platform for Spotify's advertising solutions. Let me give you a scale at how we are performing within the GenAI space. Since we launched last year, more than 70% of ads currently are using AI tools. We have around 20,000 creatives that have been generated for over 7,000 advertisers. These are real ad campaigns, real money, and real audiences. This talk is not just research. These are actual production agents that are running and delivering value. Today, we have most of our creative generation primary pipelines being generated by Spotify Ads AI. Before we dive into the architecture, this example is something that we'll be using throughout the talk for different phases of the architecture design. Let's imagine that we are creating an ad for QCon AI.

The brand itself is QCon. The message that we want to have our listeners is, let's say if you want to have production AI playbooks for senior engineers, for shipping AI solutions at scale, QCon AI is the conference to be. Our target audience is senior engineers, architects, engineering leaders. We want to target major U.S. tech hubs. Our CTA, which is the call to action, is, Register Now.

## What 'Ads AI' Actually Means

I want to make sure that we are on the same page. When I say what Ads AI actually means, imagine the advertiser says, I want to run an ad for engineering leaders across the U.S. tech hubs. This is natural language. What we do is extract this natural language into specific intent for audience. Over here, we are talking about the audience to be resolved to technology, and the geo-targeting should be for U.S. The goal is for awareness. We are using this LLM gateway using Vertex AI and our internal GCP platform. When it comes to the extracted intent, it is then provided to our multi-agent orchestration layer, which runs on Google ADK. We are using Google ADK Java. We have parallel agents that are composed per responsibility. For the ad that you just heard, we have an Ad Script Generation Agent. We also want to make sure that the ad conforms to our policies, so we have an Ad Guardrail Agent. We have an Audience Recommendation Agent that we have currently running in pilot. A combined output of these agents eventually results in two discrete objects, which we have the audience recommended and the creative that is created.

## One Agent, One Package, One Owner

Before diving into the patterns, what I wanted to show you guys was how we built this architecture in the first place. When I talk about agent, I'm talking about one agent, one package, and one owner. We have explicitly enforced this requirement because when you have the agents composed as individual packages, that package itself is composed into different aspects around the ownership, the responsibilities, the monitoring, and the prompt itself. If you see in the agent ad script folder, we have the Bazel package, which involves the dependency management of who is allowed to access this agent. We have the lib-info. The lib-info is something that determines ownership with regards to what team is owning that agent. We have an internal component that resolves the PR reviews. Let's say if you are changing a particular attribute within this particular agent, that is the ownership change that the team will go to.

Monitoring-info is basically a YAML file that has information about your dashboards and your PagerDuty alerting. You have all of this provided by our platform where before even writing a single line of code for the agent instructions, you have metrics, you have alerting, all in place. AgentFactory, Adapter, Service, and AgentModule are dependency injection scaffolding that provide a holistic package for the agent itself. Note that the LLM configuration for the agent is stored as a front matter. This is what makes the architecture a little more interesting. When I say different teams are owning different agents, we are enforcing this Bazel visibility at the horizontal layer. When I say horizontal layer, I'm saying that the ad script generation agent will not be allowed to use the audience resolver unless and until the audience resolver allows it to be imported. This enforcement is at the compile time layer. No comments.

You basically get a build failure when you are not importing that particular agent specifically. Other than the team-owned components, you have the shared platform. The shared platform provides capabilities such as metrics generation, traceability, and most of our tool usage evolves around our developer API for Spotify Ads, which calls Ads API. We also have an additional moderation layer called kutest, which provides guardrail policy management.

## Agent Architecture Stack

When you zoom out at the team level, this is what the overall tech stack looks like. We have one gRPC service. We have multiple agents. We have the runtime context, which is controlled by Google ADK. Team A owns its own agent and its own tools. Team B owns its own agents and its own tools. Team C can own its own agents and own tools. All of these agents can eventually run on a shared context and a shared runtime ADK. The tool traceability, the tool context, and the runtime loop is all managed by the platform. From an architecture point of view, here's how the flow looks like. You have the client applications making the service calls to our gRPC services, which also manage the session management. Whenever a client is making a call, it is responsible for creating a session with the agent, which eventually goes to the AI agents layer.

The AI agents layer is where individual agents actually reside. We have enforced safety and orchestration at a separate layer and not within the agent. This is because we want to have common orchestration and common guardrails to be consistent throughout the agentic experience. The last level is the model and the tool platform, which basically is the LLM runtime, the plugins and the observability configurations that are running for all of the agents. Throughout this talk, we have three attributes of this particular platform, safe content, modular agent platform, and flexible integration. These are three pillars that we want to have throughout the agentic interfaces.

## Thesis, and Playlist

This is perhaps the most important slide of the first half of the presentation. When you are dealing with multi-agent architecture, the question you always ask is, who owns what? Agents are owning the sentiment, the LLM judgment, the reasoning aspect. In our example, the use cases are like audience intent. We want to understand what the advertiser means when they say I want to target Gen Zs for our sneaker ads. We want to create the ad brief from all the signals that we get from these. Also, we want to identify sensitive topics. Let's say if a particular ad description is not conforming to Spotify's ad policies, we do want to prevent that ad to be generated. When it comes to data, we really don't want the LLMs to guess what the geo-interest would be. Within advertising, we have specific interest segments, we have specific DMA targeting. All of this data actually comes from our API endpoints.

Spotify has publicly owned web APIs that our tools are today accessing. All of this grounding happens at the tool layer. The third aspect is the application code itself. When you ask me, should everything be controlled by agents? My answer would be no. This is because whatever can be deterministically coded should always be deterministically coded. I'll be providing examples throughout this presentation where we had to establish some guardrailing and deterministic gating in order to achieve the goals that we wanted to have.

This is what today's playlist looks like. I want to start with why multi-agent at all. We don't want to just have agents because they are cool. The motivation behind having multi-agent is controlling the blast radius, having responsibilities shifted for each and every individual agent. We are then proceeding towards where to draw the agentic boundary. This is perhaps the longest discussion because it's also the most difficult to have. It is very important to understand where the boundary lies between what the agent should be doing and what should be deterministic in our code versus tools. The third is tool design. Tool design is basically prompt engineering. I say that because when we define a tool, we also define its schema. The schema has description about what the tool is doing. This description is not just documentation for developers like us, but it's also something that is going to the LLM itself and helping the LLM understand what and how to exactly call the tool. Then we proceed towards having talked about reliability through eval and tracing. In order to achieve evaluation, we need to have tracing first, and that's what we'll be stressing upon. Last but not least, I'll be talking about what we would do differently. This highlights some of the mistakes that we had done initially in our agentic journey.

## Why Multi-Agent at All?

Let's rewind 18 months back in time and think about the things that we would expect from an LLM. If we see, like last year, parsing JSON from a particular LLM model was not that straightforward. It would have been a JSON-ish response or it would have been like a Markdown with those three codes. Over time, models have gotten better. The JSON parsing, the schema checking, and business validation is now very much adopted by the model that we are using today. The third, the validation of the business logic is very important because whatever values are returned by a particular model, we can now add additional deterministic checks, whether those make sense or not. This is something which I will be talking more about in the later slides. Rewind back again to one year. If you had to write a parallel execution of agent with custom session management, you would probably have to write something by yourself.

Today, frameworks like LangChain4j, Google Cloud ADK, or CrewAI, these are frameworks that give you those primitives. The question used to be, how do we design it? Now, what should I do to implement it? When I say operations got observable, it's a supply side change. What changed in observability is the OTel GenAI conventions established some standard observability conventions for GenAI workflows. These have enabled to have generic conventions for tracing. Earlier, what used to happen was teams used to write their own annotations, teams used to write their own tracing. Now these semantic conventions allow standardized observability for GenAI workflows. With these standardized observability traces, you can now trace and replay the judge models. Replaying production traces enables us to actually test what is going on in production and we do not have to create synthetic tests. This eventually prevents us from having that black box behavior where we don't really have any idea what is going on in our agentic system.

The fourth one is interesting. It's on the demand side. The first three initiatives were more on the supply side where the technology got better. The fourth one is where the demand side increased its requirements. The multi-step pipelines have always been there in the traditional software development. What changed was, within every step, we now involve LLM interactions which can have multi-entry products. When I say multi-entry products, an advertiser can start their campaign creation on their mobile app and then continue the creation on ads manager or Ads API. Each of these entry points can have different intent and different context. Stitching this context together is only possible when it comes to a multi-agentic flow. The last one is the closed-loop measurement. When we have ads created by GenAI, we also have the ability to measure the performance of those ads. When this performance is funneled back into the system, we do know what works and what does not work.

That closed-loop measurement is something that is attributed to an overall success of an ad campaign. Multi-agents were always there. It's just that these four, three supply side and one demand side features got more better over time. We had three supply side and one demand side. Until early 2026, multi-agent was a research demo, now it is an architecture.

Let's just talk about how a campaign creation process looks like. Whenever an advertiser logs into the Spotify Ads Manager, they are providing their objectives. Like, what is my objective for my ad campaign? What is my goal? What is my KPI? What are the success metrics? This is followed by the targeting. Whom do I want to target? When, as an advertiser, I'm specifying whom I want to target, I'm specifying where I want my ad to display. What are the interest segments? Let's say I want to target technology enthusiast for QCon. Also, I want to add demographic filters. I then proceed towards the budget schedule. How much I want to spend, the flight dates, and the pacing rules. Then comes the creative uploading, which is what you guys heard in the first slide, where the advertiser provides details about what they are providing as their creative. The problem is clearly visible over here, where at the creative level, the context really dies here, because each step is isolated in its own resolve, but it does not provide the holistic signal for creating and creative.

Let me just walk over how this would look with ads manager today. We have this QCon Boston conference campaign created, where the advertiser provides their age targeting, the detail targeting has technology industry, and the languages for all the languages. I want to have this campaign for all the devices. This is an audio campaign. I want to optimize for impressions. I want to have this campaign for a daily budget of $15, let's say. Now I go to the next step. This is where I give my ad details. What is my advertiser name? What is my tagline? I should also be uploading images, which I have not. We are going to focus on the audio generation aspect of this presentation. This is exactly how the ad was created that you heard in the first slide. The advertiser provides the main message, the call to action. The call to action is something which is similar to what we provided in the slide before.

When the advertiser selects generate, they can see the script that is generated. I already have one script that is generated, but what we do allow for advertisers is to have multiple generations so that they can choose which script they want for their campaign. As you saw, while the advertiser was performing the task of creating the ad campaign, there were certain signals that were pretty much relevant in the ad creation. Today, when the advertiser is providing the inputs for creating the ad script, they are providing the brand, the message, the goal, and the ad category. These are five leaf node signals that are used for creating the campaign. On the contrary, on the right side, you can see a composable system where not just these five signals, but you can also use the resolved audience, the locale, past performance of the campaign, geo-age, gender combination, ad category, delivery-goal group. All of this can definitely provide additional context and additional information for creating a better ad. This unlock is definitely something that multi-agent systems promote.

## Where to Draw the Agent Boundary

Let's just do a comparison at a core level. At a function style, when the advertiser is performing the task of providing the brand signals, they are providing the asset and they are getting the brand tone and the call to action. When they are doing the generate script, they are providing the brand name and the CTA, as you saw in the previous demo. Over here, you don't really have the brand tone. Similarly, when you're doing the recommend targeting, you're providing the brand name and the objective, but the brand tone and content themes, they are not present. When we have an agent tech-style interface, you have the processes extracting their own signals and their own session context. After recommended targeting, all the agents have to do is just use the same shared session state, where the context accumulates across every step. At the end of the process, you have the agent getting a handle to the brand tone, the call to action, the themes, everything makes the ad creation more holistic.

This is the slide where we eventually want to go. Today we have campaign context, where the brand tone and objective are fetched from the advertiser. We also have creative outputs where scripts, visual taglines, and formats are provided from the UI display. When the campaign is submitted, it is running live. At that same time, we get the targeting information, the geo attributes, where the ad performed well, and we have the performance signals. We have the click-through rate, the completion, conversion rate, all of these signals. Imagine the closed loop actually performs as an input to our funnel, where we can say that conversation scripts with bias prompts and audience selection are actually performing better. We can tune our targeting and we can tune our model in such a way that the funnel that goes back to the system is actually giving meaningful insights about what works and what does not work. This bi-directional flow is something that multi-agent is definitely enabling us to do.

Before we got things right, let me begin with what we got wrong. As all enthusiastic AI experimenters, we started off with one single agent and a mammoth of an instruction. This agent was responsible for script generation, guardrails, audience resolution, search behavior, validation, JSON schema, and the tone and locale. It had all the tools. It had everything in one large block. Imagine something like a function with 1,000 lines of code. It worked, but it was brittle at the same time. As instructions changed, the behavior of this agent got somewhat hallucinated. We did not have an isolated approach towards identifying what is actually wrong, which is happening. This led us to determine that if you want to have dedicated responsibility for dedicated agents, this is something that we need to split as a part of our initial post POC. This is the decomposition we landed on. We are basically having three agents.

We have the audience resolver, which is in pilot right now, but what it's responsible for is resolve the geo-targeting, resolve demographics, resolve interest and exclusions. Basically, if an advertiser brief has women 18 to 34 in Nashville into fitness, this agent is resolving the interest and the geo-targeting. At the same time, we have the ad script generation, which is running two agents in parallel. We have the ad script guardrail agent, which does specific brand safety checks. We have the ad script generation agent. One key thing to note here, that this agent bifurcation is also dependent on the ownership teams. Within ads, we have different teams for different domain aspects of the ad creation and submission process. We have a team for audience management. We have a team for creative management. The creative management team owns the creative ad script generation agent and the audience targeting team owns the audience resolver agent.

This is a good heuristic to follow. A question that we always get asked is, when do you use something as an agent? When do you actually have code? If you ask me, models are getting better, but they are still non-deterministic. If the semantic of the wording of your behavior is changing, that is definitely the agent will handle or the LLM would handle. Things like senior engineers, things like technology leaders, all of this is something the agent can handle. On the contrary, if your rules change or if your catalog changes, or if your data change, that is something which is code. One good analogy is, if something that can be computed in code is computational and can be tested using unit tests, definitely have it as code. We learned it in a very interesting way. When we have targeting of audiences, on the ads manager platform, we currently support these age brackets.

We had a case where one of the inputs was women 21-plus in Nashville. In this case, the agent just emitted 21 to 99. I hope someone lives 99 years, but this was unfortunately rejected by our downstream systems because it was not part of the age bucket. The fixes that were required were pretty much obvious. We had to enumerate the seven age brackets explicitly that rejected a specific non-supporting age. The second bug was even more interesting. If a prompt says parents of teen 13 to 17, the agent actually recommended 13 to 17 instead of the parents. The fix that was required here was not the demographic, but actually the persona. These are the cases that you will find as you are testing your agents that even if you're using the latest LLM, there will be cases where you will find niche behaviors that you might find surprising that the agent gets wrong.

Another case which is interesting is the geo-targeting. A user can just say, I want to target the South and the Midwest. These are very generic terms. The South in the U.S. means 13 states. The Midwest means also 12 states. If you're not providing the country code as an additional attribute to your tool, Georgia can be resolved as a country and not the state. This is something that we learned that when you have that additional context being passed in the tool itself as a guardrail, that gives more context to the agent to resolve what specific targeting is to be provided. One more optimization that we eventually landed upon was we looked into anything that can be done before calling LLMs. Before we actually had like a multi-agent pattern for our ad script generation, we were having individual LLM calls. The problem was the guardrail was actually running irrespective of the generation.

Whenever the guardrail was giving us a false result, that was wasting token usage for us because there was no output being written for the user. We have an internal moderation classifier that actually gives us classification. What we did was we called this internal moderation classifier first. If that fails, we return early. We are doing a gatekeeping check first and only then calling the LLMs. Whenever you find cases where you might have existing solutions within your organization that can give you the output without calling an LLM, do that, because on the larger scale, you will be wasting a lot of tokens.

## Prototyped Patterns

Before we actually dive into the key takeaways for agents, I wanted to talk about the patterns that we have prototyped. This is a very common question that we get asked that, if I want to write agents, how should I write them? Or, what are the patterns that I should be using? Within our project, what we have done is Google ADK comes with a dev UI. We have scaffolded certain patterns that we currently have for our engineers, which I'm going to share with you guys today. These are primitives which are composed and modified to give different behavior based on your requirements. We have the basic router. We have a conditional router. We have the sequential pipeline, which is a very common workflow that we have. Fan-out, fan-in is where we want to have parallel execution. The last, which is the generator critic loop, where you want to refine an agent output continuously till you find the best output suitable.

Starting with the basic agent and tools. This is the simplest agentic call that you'll see, where you have a single agent and you have specific tool sets that the LLM has access to. It has zero orchestration overhead, easy to unit test and reason about. Things to watch out for is the prompt load. Eventually, you might see that 3 to 10 tool calls might be fine, but 8 to 10 is where your LLM might get hallucinated. Wherever possible, think about splitting the agent into multiple agents if you feel that you're calling too many tools at the same time. The next one is a sequential one. This is a very common pattern that we have in most of our workflows. Sequential patterns are basically a series of agent runs where one agent gets output from the other agent as a form of the output key. At the core level, you have zero glue between the steps.

These can be determined by the sub-agents for a sequential agent. The session state itself is maintained via an output key. After each agent runs, the output is provided by the output key parameter. It's very easy to add and remove stages. Things to watch out for is additive latency. The more agents that you will add in your workflow, obviously, the wait time increases. There is less scope for parallelism if you are sequentially running all these agents. Try to have the agents run in parallel if you know that these can actually run in parallel. I'll be talking about the parallel fan-out as well. One more thing to watch out for is the upstream failures can poison the entire chain. Let's say if the researcher failed in this situation, the writer will not be getting an output. It's like a single-point failure for your sequential flow.

This is an interesting agentic pattern which we have used. What you have over here is Java provides like Flowable for events. If you are extending the base agent, you can always override the runAsync function. Over here, what we have done is we have a classifier agent that classifies a query into either a billing or a technical question. If it's a billing query, it goes to the billing agent. If it's a technical query, it goes to the technical agent. We also have like a default pattern that goes to a general agent. When it comes to determining which fork to use or which route to use, this is a very interesting pattern to watch out for. The router in this case is a single point of failure. If your router itself is not able to determine what task to-dos, that is when you want to check a bifurcated approach or a backup approach for your default routing.

You do have an additional latency that adds to the router itself. Models can also over-classify intent. Let's say you have this example wherein a technical agent and a billing agent are actually doing the same things. In this case, it's very difficult for the agent itself to really determine what the intent of the user is. Parallel fan-out and gather, as the name suggests, is agents running in parallel. When you absolutely know that agent A and agent B can actually run together, these are the patterns we can use to run these agents in parallel. In our case, what we are doing with the ad script generation is we are running the ad script guardrail and the ad script generation in parallel. In this example, we have an auditor agent and a style checker. You have these two agents running in parallel, and you have the reviewer agent that reviews the eventual output.

One thing to watch out for for parallel agents is each agent does not have context of what the other agent has actually run. That is something that you should keep in mind when running parallel agents, that what was available for a sequential agent is not available for parallel agents. Obviously, token consumption is more once you have a linear scale for your parallel agents. The last one is more of a refinement approach that we have. It's called a generator-critic loop pattern. In the example, you can see that you have a SQL generator and you have a SQL critic. What the SQL generator is doing is basically generating a SQL, and the critic, what it's doing is it's providing whether that generated SQL is actually a pass or a fail. What you want to do is you want to keep on checking until you have the critic give its go ahead.

There is a possibility that the critic might just run indefinitely. What loop agents have is this attribute called maxIterations. These allow us to prevent runaway token consumption, kind of like an exit strategy. Things to watch out for is like, each iteration is a full LLM round-trip cost, so costs can multiply. Models may fail to converge without a clear exit strategy. Even in this case, you can have situations where you can run into a race condition or something like that. The critic is also an LLM. You're actually paying twice per refinement step.

## Boundary Takeaways

Going back to the takeaways. This concludes our agentic talk for what we have done with regards to agents. Agents own reasoning. You want to identify the intent, the range, the context using agents and not arithmetic. Anything that is deterministic should be code. Anything that is semantic or intent based should be agents. When it comes to validating agent outputs, schema validation is something that we currently have on each of the agents. Missing field guards, business rule validations, and null checks, those are the things that models are completely capable of doing today. Wherever possible, you want to identify the patterns for sequential versus parallel. If you have a use case where you want to conditionally exit out of your agentic runs, that is also possible today. All of this is encapsulated within a gRPC service boundary. Eventually you get your final call for your gRPC service that has the agent output, which is your holistic response. As a last takeaway, wherever possible, you should avoid an agent doing multiple things. Just like a single responsibility principle, the smaller and more concise your agents are, the better the performance is.

## Tool Design is Prompt Engineering

We'll now move towards tools. Tools are the hands and legs of the agents, that's a way to describe it. One of the things that is very underestimated is the tool schema. What is a tool schema? It's an annotation that provides the LLM how the tool should be used and what the description of that tool is. It's basically the contract for the LLM, and not just like shared documentation. The descriptions of the tool are something that are very important, and we'll talk about a few examples how these tool descriptions have actually helped us to determine geo-targeting. The way we have currently our tool setup is we treat tools and APIs differently. On a broader level, we have Ads API, which has standard web APIs, but we have agents that have different intents for calling the different tools. Now you'll understand why we have bundled agents and tools together and not as a shared entity.

If a team owns its agents, it also owns its tool. In this example, we have two tools, which are called searchGeoTargets, and a hypothetical agent that we may have in the future called lookupGeoById. The intents of these tools are completely different, but the API is the same. A key takeaway is that your tools should not be tied to just API wrappers, they should be tied to the agent intent and what the agent actually requests. Augmenting the tool with the agent requirements is more important than what API it is providing. As any other software, tools are subject to failure, and it's very important to have tools fail in a very gracious way. When I say gracious way, basically, when a tool fails, the agent should be able to understand what exactly happened. In the example for this slide, what I'm trying to show you guys is, let's say if there was a call to the interest segments, and the interest segments API, for some reason, is not providing the interest ID. There are two ways we can handle this. One is, you don't even handle that error, and you get this nice internal server error where the agent does not know what to do. Or, on the other hand, you can have a structured response, which gives the agent an idea of what exactly is happening, and exit gracefully.

Tool delegation has its different patterns and different compositions that I want to talk about. First, you can map tools as a one-to-one upstream API. This is where you want to invoke a particular API call for a particular client search. The pros of this can be like, you don't really have any overhead, you have minimal latency, and it's kind of like a direct contract. The second approach is, tools can compose multiple API calls into one single tool call. This is useful when you want to aggregate certain API behaviors in one tool functionality. Keep in mind that tool calls are also LLM calls, so these are also expensive. Whenever you are making a tool call in your trace, you'll see that that additional LLM call is also adding to your latency. There can be times where tools need a session state, or tools can actually upload the session state.

Tools have their own context called ToolContext, which can be threaded to the session scope. If a tool wants to update or read from this context, you can get that context. In our example that I mentioned, we are getting the country code from the tool context. Going back to the first slide, we have the same upstream, but different agentic needs. You can have two different agents using two different tools, but the same API. The key takeaways of the tool designing are pretty much straightforward. Whenever it comes to call bundling or splitting, you need to determine what needs to be clubbed together or what needs to go separately. Schema-level grounding is something that tools can have in the description itself to avoid static references to the data, and expensive calls can be avoided. Errors as recovery prompts, as I mentioned, are very important in terms of handling agentic tool failures and giving the agent more context around how to exit gracefully. Multi-agent wrapping is something that multiple agents can wrap the same upstream API calls with different tools.

## Reliability Through Eval + Tracing

Now that we have talked a lot about tools, let's talk about the things that can go wrong, the things that need to be tested, and the bad pattern that we initially started with. Being traditional software engineers, we love to write unit tests. These are nice functions that run and give us satisfaction of, things are running amazingly well. To boost our confidence even more, we have integration tests with mocked LLMs, where the pipeline wiring and every gRPC integration is made sure that we are wiring everything correctly. What happens to the LLM call itself? The LLM call itself is something non-deterministic, and anything that is non-deterministic is susceptible to hallucination. The approach where we say, we know it's bad when some complaints come in, that is a very bad sign, and comments like this should be avoided. When it comes to evaluation, what I want to focus on is the prerequisite for evaluation, and that is tracing.

If you have an evaluation that just verifies the output without the trace, that's mere theater. The reason I say that is the trace not only has the output of the agent, but also the journey of that output. When I say journey of that output, you have different metrics for the agent that tells you the duration, the tool call, the moderation, latency, and we also add our own custom GenAI span that records the raw input and the output. ADK offers a separate metrics plugin that we can use to bundle this behavior to each and every agent. Once we have those telemetry, we can export those telemetry to either OTel for monitoring our Grafana dashboards, or we use this third-party vendor called Braintrust that shows us a very nice demographic display of how the agent calls look. The three levels of evaluation are the tool trajectory. Tool trajectory is verifying that the agent is performing the right steps. Deterministic assets, things like schema validation, Geo ID structures, and age brackets in our case. LLM-as-judge, so identifying the tone, the brand fit, the semantic coherence is something that scoped LLM-as-judge provide a very good reason for us to have like a last evaluation step.

The last thing which I'm going to talk about is live evaluation. Live evaluation is something that is proving to be very useful. When you have an agentic run on the trace, you can actually have an evaluator which is running on live traffic. It's not a good idea to have live evaluation run on 100% traffic because that also adds to your cost. It depends on your use case. If an agent is new to the market, then definitely start with a higher value. Once you have pretty much confidence around it, then you can probably change the value. These are completely configurable values which we can change over time. One important thing that we have actually done with evaluation is we have not attached evaluation to merge builds. What we have figured out is every time you have a merge build to your repository, you are actually making LLM calls.

That is an additional cost step that we wanted to avoid. When developers are writing their own agents, they are iteratively running their evaluations offline. We do have like a nightly process that runs the evaluation for the entire suite. This is how a trace actually looks like. This is what each call and cost of the run agent for the ad script generation looks like. Over here, you can see that the two parallel agent calls are running. At the bottom, you can actually see that the ad quality scorer is running at the end. It's defining your output and determining whether the agent response is good or bad. The key eval takeaways, as you can see, are obviously, we want to have trace for every production call. We may not have live evals for every production call, but tracing is definitely something that helps. The three levels of evaluation, the trajectory, the deterministic, and the judge are the three pillars that we want to have.

We don't want to have the evaluation run as a CI gate. We want to have this on-demand and probably like a nightly build. Live evaluations definitely carry a very important signal because these are production traffic. One key thing to identify is like judge is equal to detection and not prevention. Every time you have the judge detect an anomaly, what you want to do is you want to ratchet on that anomaly and add a new assert. The more asserts you eventually have, the less expensive your judge becomes.

## Conclusion

To wrap things a bit, what we would do differently, starting with the agent architecture, is not have a mono-agent, a big, large token usage. Basically, tool design, observability, and prompt routing is something that we want to not have like a global system prompt. What I want you to take away today is you have different teams owning different agents. You have reasoning flowing through the orchestration and reasoning coming back to the funnel. That's what the whole multi-agentic architecture currently looks like.

**See more [presentations with transcripts](https://www.infoq.com/transcripts/presentations/)**
