Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products Elastic principal data scientist Susan Chang described how the company consolidated per-team agent evaluations into a shared, tracing-based framework for its production AI agents, including the attack discovery security agent and enterprise chatbots built on Elasticsearch. Chang said each team previously created its own datasets, evaluators and tracing storage in separate locations, and the shared framework is intended to estimate response quality, prevent regressions and repeatedly test agents with less manual work. Elastic's agents run on top of Elasticsearch, which Chang said one banking customer uses to ingest petabytes of data for cybersecurity. Transcript Susan Chang: My name is Susan. I'm a principal data scientist at Elastic. We make tools like Elasticsearch and Kibana. I've also written a book with O'Reilly called "Machine Learning Interviews." Today we're going to be talking about a few things. I'm going to talk a bit about what kind of agents we've been building and try to maybe connect them to what you folks might have been building. Tracing, which is, in my opinion, very important and a huge foundation, as well as how we use that to enable all these evaluations. I'm going to talk about how our company made this journey through maybe individual, like rebuilding a lot of evaluations towards building a shared framework, as well as what are the building blocks we have in that shared framework. Then tying it all together, and then see what takeaways that we have there. AI Agents We Are Running in Production Just so that we can be roughly on the same page, I'm going to describe a bit of why we're building agents at Elastic and the type of agents that we're running in production. Elastic, we make Elasticsearch, so a lot of companies, they use us for just search or retrieval. They could use us for like normal keyword search or vector search. Every time you've been searching on Stack Overflow or using Uber and whatnot, they use some sort of Elasticsearch tools. GitHub as well, and Tinder, these are just some examples. I mention these because there's just a really wide range of industries and use cases. It's really hard to describe what exactly people do. I'm trying to put some examples as well as what we're doing internally. What is the bottom line here? People just have data in Elasticsearch and then they can apply it to different use cases. Internally, we actually build on top of Elasticsearch, observability as well as security. This is important because I'll be talking about some examples where we actually, in Elastic, we build reusable security analyst tools that use AI agents on top of Elastic. This is one example. The use case for this is that like an Elastic user, they could already be storing a bunch of logs in Elastic. We have one huge banking customer. They said that they're ingesting petabytes of data into Elastic for the cybersecurity use case. Why is this important? When they have a cybersecurity incident, they need to be able to look through all those logs, filter them, and extract the right information. They already have all this data within Elastic as a datastore. In Elastic, too, we want to build some reusable things for the user because they might not be power users. They want to be able to extract the logs or just use AI to read the relevant logs and then summarize it. One thing we've built in Elastic is called attack discovery. It's like, I'm using agents to be able to pull from logs and then find possible attacks instead of a human going through and manually parsing through them. There's another agent that we've built that is also on top of Elasticsearch where people build enterprise chatbots using their proprietary data that's stored in Elastic. I mention this because these are super different examples, but because people use Elastic for so many different things that we also build them internally so that people can reuse them even if they're not experts. When we're building these agents, we want to estimate the quality of the responses and prevent regressions when we make changes, as well as repeatedly test them without too much manual work. This is what we would build for each of those. For that attack discovery agent suite as well as for the enterprise chatbot situation, each team would create their own datasets, they would create their own evaluators. Then you would trace the data in different places, we'd store them in different places, and then do evaluations in completely different locations. Just as an example of what we would build, this is that cybersecurity example where we have AI agents that would extract potential cyber-attacks. This is what we would do. We collaborated with security experts, like security analysts and researchers to find some novel attack scenarios and they would create these scenarios, which are for evaluation purposes. The AI agent needs to be able to identify these attacks if they happen. We would need to create these datasets. We would need to create and store all these benign scenarios as well. We would need to use these metrics. This talk doesn't really cover all the metrics in very huge detail, but I do want to show you this example because the security agent builder team, or the team that's building the security AI agents, we are using precision and recall to match alert IDs. We want to prevent this attack discovery AI from hallucinating. We also use similarity scores, factuality scores. MITRE tactics I mentioned here are like a security domain, like taxonomy, classifying different types of cyber-attacks. We want to make sure that the LLMs are not hallucinating unrelated MITRE tactics, for example. We also have some domain-specific metrics as well, because the security agent team wants to make sure that our agents are not performing poorly. For this other team that's building the enterprise chatbots, they're creating a completely different set of test scenarios. We have different indices using like query and retrieval scenarios. I think I have an example here. This is just like one example, which is, I'm trying to verify my domain with Google Workspace. Do I need to add this to my DNS and whatnot? It should be extracting and answering based off of existing documentation. This is what their data looks like, completely different from the security AI agents team. This is more for like a chatbot team. They're also using precision and recall. They're using factuality. They have some other metrics that they're quite concerned about, the response relevance, response completeness. This is pretty interesting too, ES|QL. This is like a query language that Elastic has built. We have actually maybe a bit too many query languages. We have like KQL, ES|QL, Painless script, but either way, for AI agents that are built to answer things about Elastic, they should be generating correct syntax. They should be generating correct ES|QL. We also were able to build into those AI agents calling tools, to be able to format or check these queries. This is important. This team is concerned with that, whereas the security AI agent team was not so concerned with that. We have even more teams that are building different types of AI agents. I think for folks, your company might be in various stages. We've been building a lot of different use cases for maybe a year, two years or more now. We have other security domain AI agents, and then observability as well, and so on. As we build out all of these different agents, or skills, or components, we want to be able to evaluate how they're doing more and more needed. Each team is building their own agents as well as evaluations. Tracing Your Application Just to go through this part, I do want to emphasize that this is pretty important for tracing. There's a lot of tools. We have heavily used LangSmith and Phoenix as well within Elastic, because Elastic already has the capability of observability for a long time. I think that's not the only tool. We have all these other tools that we have tried out. Other companies might be using their own. We tried different tools. I think because the industry is moving very quickly, it was really important for us to see what fits our use case. Then we're also building things and learning as we go. We wanted to trial out different tools as well. This is an example. Let's say this is the design of your AI agent using LangGraph. This might be how it roughly looks like from a conceptual level. When you trace that, you want to be able to see all the calls. Let's say it's an enterprise chatbot case. Actually, when you type, can you help me resolve this problem with an Elasticsearch index? It needs to run vector search, run keyword search, and the AI agent's doing that under the hood. Then you need to know this is what it decided to do in response to this user query. It has very granular amounts of information of what your agent is doing and what tools are being called, what vector search is being run, what databases and data is being retrieved, and so on. One other example using the Elastic Observability, it's pretty much the same, or are using Phoenix. Either way, either tool, you want to make sure that you're capturing the data that you need for evaluation, because let's say you've built this AI agent, and you're only evaluating it on the end result. You don't know what happened in the middle. You don't know that the agent is actually querying the wrong database. You just saw that it has the wrong final number. It's really hard to actually take action on these evaluations if you don't have all this information. Or, with this information, you could also run evaluations on specific tool calls, things like that. One thing I really liked about LangSmith is this Add to Dataset feature, because, as new scenarios come in, you might have a situation where the AI agent actually made a wrong response. Your user might have given you some feedback. Let's say they gave you a thumbs down. Then you track that. You're able to add that to a test dataset that you can run on future iterations of the evaluation, to make sure that your AI agent now performs well on these situations where it used to not perform well on. One anecdote, because I've heard people ask, why do we need to capture so much information? As well as, of course, there are industries where you cannot capture all of that. That's fair. I've worked in more restricted industries as well. We do have to make use of a lot of proxy signals. In terms of other types of machine learning use cases and evaluation, we've always had some sort of telemetry or a clickstream data if you've worked in recommender systems or whatnot. Even if you're unable to capture all those signals, we've always had to make use of some sort of proxy signal or explicit feedback, implicit feedback, and so on. I think this is very similar to that. Shared Evaluation Framework: Overview Now to get right into this shared evaluation framework. As I mentioned, it is really cumbersome to keep building the evaluations from scratch for each team. We wanted the different developer teams to collaborate and maintain a shared framework and make it really quick to spin up fully featured evaluations. One of the examples after building the shared framework was that we created AI rules generation. Rules are like YAML/JSON configuration files. They follow a specific structure. Then we made this agent so that users who are not very knowledgeable about how to create them, they can just type in natural language, create me a detection rule that can flag users that logged in on the weekend, or things like that, like anomalous behavior. I was able to spin up the evaluations for that and focus more on other aspects of the work instead of rewriting an entire evaluation suite from scratch. This does take some time. Initially, the security organization and the agent builder/chatbot example, built out completely different evaluation frameworks. Then we had to go through some consolidation there. This is what the shared toolkits, they do. It can import the datasets. I showed earlier for the cybersecurity case, they have completely different types of data than the data that the chatbot case uses. The chatbot case uses like Q&A type. Whereas for the cybersecurity, it takes in security logs and then needs to extract potential attack scenarios. It needs to import different types of datasets. We do have a shared and standardized schema. We also have the trace-based evaluators, which are more related to like observability side, like token usage, latency, tool calls, performance. We also have shared RAG evaluators. I showed earlier some examples. We're building the enterprise chatbot example, but they also have different types of AI agents that are using RAG under the hood as well. As well as any custom evaluators that are needed. This is one example of how one evaluation run would go, which is, we have the datasets, which have the input and output. It would call the task or just make an agent run. In the enterprise chatbot scenario, you would have a question and an answer. We would run the evaluators, the code-based ones, LLM-as-a-judge, and other trace-based ones. These are customized per use case. Then store the results back into Elastic or other types of tools that I mentioned earlier. This is like a high-level workflow of what is reused and then what can be customized within these building blocks. As a data scientist, this part was really interesting to me. Just as an aside, a lot of the codebase is written in TypeScript. As a data scientist, we originally used Python to write these evaluations. We used Python to load the datasets. We used Python to call the AI agents or run the equivalent of the AI agents, which were actually in production in TypeScript. We had a little bit of a mismatch between what's actually in production and what we're comfortable with. It was interesting because originally in the codebase in TypeScript, they were already using Playwright, which is a framework for running tests and APIs. We worked together with these software engineers, as well as using a lot of Claude and Cursor and whatnot, so that we're able to translate a lot of those tooling from Python to TypeScript, so that we could more closely match what is actually in production. This was definitely interesting. I think it's like an example of potential AI enablement as well, which is that we are stepping outside of our comfort zone with using Python. Actually, in TypeScript, there's a lot of stuff that is not really supported by a lot of the mature frameworks that Python has, but just that, for this scenario, we wanted to match production more closely. I wouldn't say that it's applicable for every single case. I'm not saying, go and translate everything, port everything to TypeScript from Python. No, that's definitely not what I'm saying. Just like for the production case, we wanted to do that. We are running on Playwright, basically, and with a customized Playwright, which we call Scout. This runs everything. It loads the datasets. It also runs the AI agents, which are implemented in TypeScript. Then it does all the collection and then the evaluations. That's just one thing that I wanted to talk about. It's like related, but it's more about our experience as we develop to these shared tools. Also, from the developer standpoint, because in the codebase, like we're using TypeScript, so we use a lot of local development as well. It can also run everything locally and then output the scores and results on the terminal, which is convenient from a developer user experience. Shared Evaluation Framework: Building Blocks These are the general building blocks, next on, I'd like to talk about. I mentioned the high-level workflow, which is that we saw that different teams were using these metrics like precision and recall. I showed that in the security use case that we use those. We also tested against factuality, and so did the enterprise chatbot team. They also were using these metrics. This is the type of metrics that we eventually consolidated between teams. Semantic similarity, probably mentioned a bunch in other talks. Actually, I heard some mention of this in a previous talk as well for a semantic search, where you want to identify if what the semantic search or the vector search pulling is actually the most relevant documents or the most relevant information that's being pulled by the AI agent. This is all important for different types of AI agents that are making use of RAG or some sort of search and retrieval. The third one, factuality, which checks against like, let's say, product IDs. If you're building a chatbot example, you don't want to hallucinate a product ID that's not there. For the cybersecurity use case, we didn't want it to hallucinate MITRE tactics. Those are categories of cyber-attacks that don't exist. We also built some shared tooling for LLM-as-judge, or you might call it model-based evaluations in the shared tooling, because we found that a lot of different cases were using LLM-as-judge, but just we had implemented it separately in a different way. We actually did start out with all these ad hoc evaluations and then figured out what we could consolidate because everything was moving pretty quickly. I think if we had sat down for two months and planned this out and then built it out and then started building our agents, I don't think we would have been able to ship so much. It depends on what stage that your organization or that your company is at. This is an example of how we did this in Python, but as I mentioned, we later did move it into shared evals in TypeScript. I have showed this example elsewhere, for more like a data science audience. I wanted to say that if you're familiar with Python, you can start building these evals right away, even if you're not familiar with what the actual production stack is. You can still build them. You can still build your evals. Later, if you're thinking there's a mismatch between your evals and what the scenario is in production, then you can later merge them. This is one of the examples of what we did in the data science team. LLM-as-a-judge, just as an example of the ES|QL, as I mentioned, that's like an elastic query language, a user might ask something like, generate an ES|QL query that can do this or that, like search data. In other cases, it might be like, I want to use generate SQL, like SQL or something. We at Elastic, we own ES|QL, like this query language. We have to make sure that it performs well in our own AI agents where the user is asking us to generate ES|QL queries. We would have known good answers in the test dataset. The shared evaluation framework needs to be able to pull these known examples and compare them over the test dataset. A very simple LLM-as-a-judge is just, are these responses similar? Of course, as things mature, it became much more complex and more nuanced. This is just like a starter example here. Some of the pros and cons of the LLM-as-judge that we encountered, but we eventually, of course, still decided to put this in the shared evaluation framework, is that we definitely used LLM-as-a-judge a lot. I think at the beginning, we might have this scenario, if you use LLMs to evaluate the performance of another LLM's response, is that just like garbage in, garbage out, or like evaluating garbage and whatnot. We did actually try it for a lot of different scenarios, and we found that it is useful. As long as you are pairing it with the right type of hybrid evaluations, it's good for scaling to different ambiguous scenarios. Or let's say you have an AI agent that is meant to read in a certain tone or brand tone, coherence or style, these are all pretty good situations where the LLM-as-judge probably performs better than other types of evaluations that I'll mention later. It's good for open-ended tasks, as well as you can self-grade it using chain-of-thought to explain why it came up with the results. Like, it didn't fit our brand style, or it did fit our brand style and tone, and why. This is one thing that we found. The cons are that it might not be granular enough. I mentioned a few scenarios where in the result generation, it might have stuff like query language, output, as well as different things, or JSON objects, or YAML, where it needs to follow a specific format. The LLM-as-judge might not always perform best by itself, so we have to combine it with other tool calls and things like that within the evaluator. Also, another issue would be that there is different maybe results across runs, even with the same model. LLMs might not be able to identify things like internal info, like product ID. You'd need to combine that with other things like tool calls or more like programmatic evals, and so on. This is one interesting thing that we found as well. This is for a cybersecurity use case, which is why, as a team, we were researching these cybersecurity evals, and Meta also found, and CrowdStrike, which is that if we're using the same family of models, in this case, like Llama to evaluate Llama, then they would think the Llama models perform better. This is something that you might want to pay attention to if that causes any unexpected results. The interesting thing is like, we have been doing this for maybe a year or two years, and we had a bunch of this stuff written up internally, and this January, Anthropic released a blog, which covered a lot of the things that I had mentioned and that we had found as well in reality. It was a pretty good validation with regards to their opinion on LLM-as-judge and model-based graders. Next, we have what we call rules-based evaluations or programmatic evaluations. Other companies or other teams might call it different things, but I'm just going to use this for now, which is catching the crucial points that LLM-as-a-judge is unable to do well. Using code like Python, TypeScript to check if generated code works or does not work. Heuristics, let's say you can extract different entities. Let's say, for this chatbot example, it needs to output the correct product ID. If it didn't output any product ID, then it should be wrong. Just simple scoring logic that you would want to have that you do not need to rely on LLM-as-a-judge for, as well as programmatic syntax checks, so on and so forth. Things that you do not want hallucinations for, things that have a deterministic answer, if you're able to break that out, then, in our experience, that really helps with evaluating the outputs of your AI agents, especially if they're working under some ambiguous circumstances. Combining it with the LLM-as-judge is ideal. The pros and cons here of these building blocks are that there's no ambiguity. It catches some of the weakness of the LLM-as-judge. It's cheap. It's fast. You don't have to make an LLM call for just every random evaluation. It's harder to scale, though, for larger problem spaces, as well as natural language. It's really hard to use a rules-based or programmatic evaluation on a full chunk of text and say, this answer is correct. It's better to let LLM-as-judge handle if it's a larger chunk of text, natural language. It's a bit weaker at open-ended tasks, which is where we try to use them both together. What Couldn't We Abstract? In the end, what could we not abstract into our shared evaluation framework? I mentioned a few of these building blocks. It does overlap a bit with what each individual team has done, but I did want to mention it here because those were the things that we ended up abstracting out, because every team was using all of those things in very similar ways. That made sense for us to abstract out. What couldn't we abstract out and what each team should be owning themselves? We could not really abstract bespoke data creation. I mentioned in the cybersecurity use case, in the security use case that is still required, we had security analysts and security researchers that are giving us a user perspective. If you have a cybersecurity analyst AI agent that's supposed to help security analysts find attacks faster, we need to know from them, not from software engineers and data scientists, if this is working well. We need to know from the actual end user what the test dataset should be looking like. That is pretty important, like having domain expertise to contribute to the dataset creation. That's not something we could just automate away in a shared evaluation framework. Also, on a related note, product input. If you work closely with the product team in your company, what are the positive examples, what are positive behaviors or negative behaviors of your AI agents? What are situations where the AI agent could get things wrong a little bit? You have to add them all into the scoring mechanisms for the evaluations. Regression might mean a different thing for different agents as well. Sometimes we have released updates and then it might have previously performed well. Now it became more aggressive and didn't perform as well. One of the examples for the security agent example, it used to not hallucinate some attacks. After we released an update, even if you gave it all benign data, it started to think, yes, there's someone attacking our environment. We were able to fix the evaluations to catch those behaviors. A lot of this is iterative. You might not know exactly what counts as positive or negative examples or actual regressions. It's hard to predict everything, especially in this more ambiguous workflow. The shared evaluation framework cannot be responsible for calibration. Calibration, basically that if you're running the evaluations, they should roughly be performing the same. Your evaluators generally agree with your human evaluators or your target behavior. If the evaluators are not well calibrated, then what it's going to give you is just junk. Let's say you're using LLM-as-judge and then you run the same scenario three times, and the LLM-as-judge thinks it's a different performance each time. Then that means it's not well calibrated. We had to further improve these things. This is all reliant on domain expertise, as well as each team's individual expertise. This is not something that we have in the shared evaluation framework, because we can't. Key Takeaways Some of the key takeaways. If you're building the first agents in your company, it's definitely a little bit different for each company. This is how we started when we were building our very first agents, very first LLM-driven use cases. We were working pretty closely with the domain experts. I think one thing I really remember is that the cybersecurity cases, we worked together with the analysts and they would spin up VMs. They would spin up different Okta, like test environments and things like that. Those are pretty hard things to do. If your AI agent is meant to do something really complicated with that, don't skip the product and domain expertise. It is tempting to bypass that and just ask an LLM to generate a bunch of stuff, which we eventually did to scale up the test dataset. We made sure that we anchored that to a lot of this real expertise and real user needs and things like that. In this early stage, when we were building our first agents and other teams were building their first agents as well, we tried out different tracing tools and see what works. As I mentioned, we actually started building agents pretty early on with LangGraph. It made natural sense to start using LangSmith at the time. Other teams, they started using Phoenix or whatnot, and other teams still you might use Elastic Observability. This is what was happening early on. Starting small is also something to note when starting out. I have spoken to a lot of people from maybe folks in this kind of stage, and they're a little overwhelmed. How do I create this dataset? It sounds so complicated. I'm making an enterprise chatbot, how do I gather these question-and-answer examples? I've seen from people in industry, and actually also from Eugene Yan's blog as well. He currently works at Anthropic. He also mentioned that, for some situations, just starting with a sample of 20 records for your test dataset, to 50, and then you can later scale it up. I think it is important to answer, especially if you're at your first stage, and then that means the AI adoption or AI development at this company or organization is more so starting out or nascent, then that means ad hoc work is fine. There's nothing wrong with just having a small amount of data. It is important to evaluate. That's how you can justify further building out of these AI agents in your company or even scaling up to different teams and different domains. I think it is fine to not overinvest in abstractions at this stage, but then optimize for more learning and experimentation. That's what we did. That's our experience with these different teams. I think, actually, early on, we might have worked in a bit of a siloed situation. Like, wait, you're using Phoenix? Wait, we're using LangSmith. Something like that. That might happen. That's ok. I think I heard the speaker from LinkedIn, he was saying that some of it is an organizational thing. Like, if you have a shared way that your company is using LLMs, that is also a challenge and that is important. I echo that. Communicating cross-teams and seeing what you're aiming to do. If you have a few agents in production, so this was us when we had maybe some agents and then we had different evaluations. We did have them, but if you don't happen to, definitely start building at least one. Product and company leadership is going to start having questions about how your agents are performing. Users might start having troubles in production and you want to be able to troubleshoot them. This is why you need to at least start building some tracing or evals. Thinking about the common metrics you're observing and iterating on that. As I mentioned, for the security use case, after we pushed some changes, we realized that they were hallucinating attacks on benign data. That means that we had to go and refine the metrics. It's all iterative. You might eventually run into the problems of too many different evaluations maybe caused by siloing or caused by different product teams building their own thing. Also, wanting to move fast. I think that's fine, too. That's what we encountered as well, which is like, you do have to move fast. You might have to just stick to one tool and decide. Consolidate later. As the agent development matures, then all the shared tooling should be consolidated further on to avoid a lot of pain. This talk is about the evaluation framework. We also had done a lot of work to consolidate all the agents as well. We used to have security agents that does this. Then I found that I had to click a dropdown to switch to another agent that was also on Elastic. Then in a recent release we consolidated that workflow, and it can identify what agents you're talking to instead of going to a dropdown just because you clicked onto something else. Those are things that we also consolidated. This talk is not about that. We did have to consolidate agents as well as the agent evals. Lessons Learned Lessons that we have learned, if we could do things over. Building first and then consolidating later. I think if I had to do it over, we would start communicating a bit earlier, and also communicate with the goal to work together and consolidate. Because I think for some time we're like, your team does that, our team does this. Maybe can you give me some of the access to your platform? Then they would just add us to their team on Okta or something, ad hoc. That did work. Then we weren't really talking about consolidation at that time. I think that lasted for some time. If I had to do it all over again, I would really try to do that. I would also try to make it easier for feedback to improve the product. I mentioned proxy signals, and things like that. Because for a lot of the enterprise users for Elastic, we can't trace all their data at such granularity. They have to agree to share it with us. Even if they share it with us, it's pretty redacted or pretty bare bones. I think if we're able to build a better way for users to provide feedback, this might be similar to thumbs up and thumbs down, you might see on ChatGPT or things like that. Whatever mechanism, even if you're only getting these proxy or indirect signals, at least having some feedback mechanism or a structured way to track. For another machine learning product we had, we use Google Forms. At least you have something. Have something like that. This is a really interesting thing, like the main stack. What I mean by that is that in our situation, because everything was in production in TypeScript, and we had built and invested a good amount of time in building the Python evaluation framework, if I had to do it over, I'm not sure if I would exactly not do the Python one at all, because that's what my team was specialized in. I'm a little bit half and half on this one, because I do think when we did jump into using TypeScript, and we were using of course Claude and Cursor heavily for that, with the review of software engineers that are more experienced in TypeScript, that did enable our team quite a lot to directly contribute to the shared evaluation frameworks. I do think this was a good decision in the long run. I don't think we would have started out that way, even if I had to do it again. If you already have your own shared evaluations, what would you do over? I think that's something that I'd be curious to know. We do share some of these evaluation results publicly, not for every scenario. I shared at the very beginning of the talk that different industries and companies use Elastic for different things, and also internally, we try to build these reusable building blocks for our customers and users in different domains. We do share some of our recommendations for users in security domain and whatnot with evaluations like that, but not every single thing. This is interesting having a bit of public scrutiny or sharing our work a little bit, which does demonstrate that it is nice and useful to have evals so that you can see what you might recommend to other teams in your organization or even to the broader public. Questions and Answers Participant 1: We currently use Phoenix in-house for the tracing aspect. I'm curious how you think about storing the data in Phoenix versus in an application database or something else. Is all your eval data in the tracing framework, or is some of it in the application database? That's the first question. Then the second as a follow-up of tracing, if you compare it to OTel and now with the rise of harnesses and deterministic steps, sometimes the failure is not just in the inference steps. Let's say, it might be a deterministic step that is a failure, how do you think about incorporating that into the overall debugging and training process? Susan Chang: The first part was around Phoenix and where we store the data. This is a bit of a hard thing, because in some organizations and some companies, you already store observability data somewhere, so that makes it easier to make a decision. In our case, it actually went like this. I believe we had our observability data stored somewhere else. Because that observability, including our own, didn't have a lot of the LLM-specific tracing out of the box at the time, we just decided to get something like Phoenix that had built a lot of that, the useful add-ons, for specific LLMs or agents tracing. We had this data stored in other places. Let's say you already had a lot of observability data stored somewhere, then I might as well just use the same thing, if I had to start over and if every tool was mature already. That's the way we did our decision-making, which is, yes, we did store them in maybe different places depending on what we wanted to do. All of our applications already had observability that we were using for a very long time that was in some other thing. That's one part. I think I've heard people, let's say they use MLflow, actually, for their agent. That's because they had already had MLflow for some other machine learning use case and whatnot, so then they ended up using that, even though I'm not sure if it was because it's the most mature or because the data storage was actually easy or anything. It depends on what's easy and what your company already has, and whether your company has maybe an appetite to try out different tools. Because I think at the time, a lot of those tools are not that mature yet. This was since a year and a half or two years ago. Some of those tools they were building as we were trying them out, and they changed so much. Participant 1: As agents stop becoming a one-shot thing where you just provide context into a black box, as there's more harnessing and deterministic steps around it, how do you think about, from a debuggability and tracing aspect of incorporating the deterministic steps as well into it? Because sometimes the failure is not in the inference, it's in the deterministic steps. Susan Chang: For us, when possible, we do trace everything in our developer environments. That does include a lot of those deterministic steps and things like that. When we do evaluation, it's not only on the end result, but sometimes we might do it on, is it making the right tool call, for example. We do have a lot of the granular evaluations as well, depending on what steps in the middle the agent might be taking. That's only because we have that data available that we're able to evaluate those things. Any deterministic thing, if it can be evaluated and we have the data, we are definitely taking that into account. Participant 2: I'm wondering if you could go a little bit into, once you run the evals, what kinds of considerations go into setting the criteria of, yes, we're going to promote this agent, or, no, we're not? Are some of those criteria either pass, fail, or is there some way of setting lower bars? Or are you just looking for improvement? Just looking for some general thoughts on that. Susan Chang: We have a matrix that just has the granular scores, like let's say precision, recall, and then factuality, and other things. Or maybe if we're evaluating specific tool calls or things like that, we'll have specific metrics. Then we'll have a roll-up score, just a simple weights thing. Let's say we think factuality is really important, then we'll weigh that higher. Then we just add them together and then have a composite score that's flattened into one single value. We also have the other scores that are available so people can see. Because when you flatten it, it's really easy for decision-makers or even for ourselves, like we just want to see one value at the end. We might have a threshold and it has to improve, but it needs to be over a certain threshold as well. We also want to investigate the maybe more detailed granular metrics as well if we can. I think the production is actually more simple. We just flatten it into one score, threshold, and then we might run some ad hoc tests, manual. Manual kind of like use the product and click around and stuff like that, too, to do a last smell check. Then we release it. It's rigorous, but also not really. See more presentations with transcripts https://www.infoq.com/transcripts/presentations/