# Presentation: Ontology‐Driven Observability: Building the E2E Knowledge Graph at Netflix Scale

> Source: <https://www.infoq.com/presentations/netflix-observability-aiops-ontology-scale/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global>
> Published: 2026-10-09 11:00:00+00:00

## Transcript

**Prasanna Vijayanathan:** My name is Prasanna Vijayanathan. I'm part of the observability team at Netflix. I work on observability, performance, and quality of experience across all our member applications. What that means is from the point you click on the Netflix logo until you start watching something, how does your experience feel on Netflix is what I look into. What metrics should we look at? What tools should we have? What insights should we provide to our users to make sure that we understand what you all like? If you do like watching Netflix, you're welcome. If you don't, you can blame it on my colleague, Renzo.

**Renzo Sanchez-Silva:** I'm Renzo Sanchez-Silva. I've been at Netflix for almost 10 years now with the monitoring and alerting team within observability, building several alerting infrastructure, and now getting into AIOps space. I've done also some research back in the day in ontology. With AI, everything is converging.

## Observability at Scale - Why is it Hard?

**Prasanna Vijayanathan:** Today we're going to talk to you about a problem that we've been tackling, which at the surface feels like an observability challenge, but it's really a hard data engineering problem. Before I get into that, let me set the stage up. Why is observability hard at scale? What scale are we even talking about? Let's do a pop-up quiz. Netflix launched live events. End of 2024 was one of our first major big live events, a boxing match between Jake Paul and Mike Tyson. A few of you might have watched it. What do you think was our peak traffic during that event? It's definitely in the millions. It's 65 million concurrent streams. What that means is there are 65 million users clicking watch and watching the stream at the same time. That was our peak traffic. One more. When you log into Netflix, the apps are sending requests into our infrastructure for our backends to respond and stuff.

How many requests do you think we handle per day coming from our applications to our servers? We handle more than 2 billion requests a day. Last one. With all those requests happening, we obviously have our own observability systems where we log metrics, a lot of different things. How many real-time logging events do you think we have to handle per second? It's upward of 38 million events per second. That is the scale we're talking about. Numerous client platforms and devices across all the TVs all over the world, mobile phones, laptops, you name it. Thousands of microservices emitting all these billions of telemetry events. Every single member touches so many different parts of our infrastructure, our CDNs, nodes, and so on. On top of all that, we have thousands of experiments running through A/B tests throughout our ecosystem. Managing all of that at scale. It's everything, everywhere, all at once.

## End-to-End Observability - What is It?

Even at that scale, our goal at Netflix is to deliver the best possible experience to every user every single time. Our strategy towards observability is simple. Do we have the right tools, metrics, and insights to enable us to do that? The strategy is simple. It is not easy. At this scale, correlating user experience with system performance is a fundamentally different challenge. It is not a tooling challenge anymore. It is a data engineering challenge. Traditionally, observability systems have been monitoring and alerting tools where you observe your apps and services. Something goes wrong, you get flagged. Then you look at what's going on, triage it to the right teams, look for the right developers to fix it if needed, roll out a fix, and then see if the data comes back to where it is expected to be. That's how traditionally observability systems have worked. With end-to-end observability, we wanted to rethink how we do observability, and change it from a reactive monitoring experience to a proactive insights engine.

Before I show you how we do that, I want to talk about the vision that we created. Imagine a system that can detect issues automatically in a few minutes across the entire stack. It's not just one service or app, but across your stack. Wherever an issue happens, you know that this is what is happening. Then you go into automatically prioritizing those issues based on how it's impacting our users. You also can triage it to the right team automatically. You don't have to look at one app and one error and say, I think I'm going to follow this needle through this haystack and pick out different services and loop them in. Then, can we automatically root cause? Not just triage, but can we actually identify the root cause of the issue? Or, even better, can we predict issues even before they would actually start impacting our users and maybe even start suggesting or taking corrective actions?

Is this possible? We wanted to do this. Building the system would give us seamless visibility from the user's device through our networks into our gateway services into the deep trenches of our backend dependencies, services. All connected, all in real time. Users would have no disruptions to their experience on Netflix. Every time you watch Netflix, you will love it. Your experience is going to be awesome, every single time. This is our vision.

Sounds great. That's far from reality. It's very different today. I'll walk you through a timeline of a recent investigation, probably within the last year or so. There was an alert that got triggered on one of our client platforms, I'd say around 3:30. Typically, on-calls look at it and they're like, let me wait for a while. Let me see if this is going to sustain. Maybe something happens magically and it'll go away on its own. Sometimes it does. Sometimes it doesn't. A couple of hours later, they're actually looking at it, "I think it's a real issue. I need to go figure out, follow this needle and see if there's another service that's impacting." In the average case, usually it does. Then it's triaged to the right team. Then it turns out that team was already debugging this issue from like 1:30 that day. People are already on it.

Then, multiple teams are involved now. There's a lot more debugging happening between all these different teams. Then once the debugging starts, they root cause the actual issue, follow the needle down to the actual offending service. In this case, it turned out there was some service that thought they're not as important and cut down their capacities. That caused backward propagation errors. That's around like 6:30, three hours since the client application actually noticed real user impact. We rolled out a fix. It was resolved within the hour. This is an average case. It doesn't sound too bad. It's like four hours. Not too bad. Then when you consider what it took, during this time, there are nine teams that were paged. More than 30 engineers were involved in the entire debugging process in different capacities. Then three other related incidents that were also flagged. There's panic everywhere across teams. This is typically, and this is an average case. This is not even a bad example. Then when you put together the four hours and all the resources that get involved, it just feels wasteful. This is what you want to solve with end-to-end observability.

We looked at why. Why does this happen? We came up with four main themes that existed. One, you've got too many siloed data sources. We got roughly logs, metrics, events, and traces that are stored in different forms in different datastores. They are in silos. They don't talk to each other. They're not standardized. Structures are different, and so on. That leads to disconnected and non-contextual alerting, which means each team follows a data source that works well for them in their silo and then set up alerts. Then when alerts trigger, they are looking at the same problem in one way compared to another team that's looking at the same problem from a different perspective. They should be talking to each other, but that doesn't happen. What that leads to is complexity in how teams triage and actually troubleshoot the issue. Ultimately, what that means is we actually don't detect what is the impact for our users. If you're a downstream service, 5 hops, 10 hops away from the user, you almost never understand how your system and service is affecting the end user.

## Connectedness - Bridging Gaps, Breaking Silos

Given the vision that we saw, that we have for end-to-end observability, how do we solve this and move from the state of reality to the vision? The one-word answer for that is connectedness. That's what we thought. The common theme is that. We want connected data, connected alerting, connected debugging systems, workflows, and root cause analysis. If we create that theme for end-to-end observability, we don't need 10 different data sources giving rise to 6 different tools. We want to create an integrated observability layer. One abstraction that can tell you everything about our system's observability. What that means, then, is our dashboards, alerts, debugging workflows, and incident response tools, they all speak the same language. Then duplicating efforts in debugging and triaging, the same issues, and the same symptoms across the system is reduced, which means you can follow a breadcrumb trail, and then get into the same path from the user's device to the actual root cause.

You can follow the breadcrumb from different trails, but get on the same main path of root causing the issue. That leads to reducing the time it takes to resolve issues and ultimately will improve the accuracy in diagnostics, and give our users a less disrupted and better experience. The core piece of all this starts with the data, connected data. We got our users, our devices and applications, and our services and infrastructure, all emitting these broadly four different categories of data: metrics, events, logs, and traces. We unify them and put them in one layer, the MELT layer. Once we unify all these sources, we solve the problem. It's important that even if all of this exists together, we need them to be accurate and actionable.

Let me walk you through another scenario. Let's say we have this layer already. Let's say there is a regression we see, LOLOMO TTR. LOLOMO is the metrics of the list-of-list-of-movies that you see on Netflix. TTR, some rendering time for that list-of-list-of-movies. You see a regression there. There's a slow leak, and there's a big regression. Developer notices this and they ask our unified data layer, what happened? Now all the observability data lives in one layer, so tell me what happened. Then the metrics system would tell you, yes, I see LOLOMO TTR regressing. I also see 12 other metrics across 5 different services correlating, happening at the same time. They are also regressing. We find some correlations. That's good. Then events could tell you there are 47 events across all these different apps and services that also happened right before this regression started. Yes, I can see it on a graph.

I can see all these events along with the metrics. Then logs tells me, I think there is a connection pool exhaustion error that spiked up at this time in one of those 7 services. Maybe that's what it is. There are also all these timeout errors in all these other services at the same time too. Then traces can tell you a different story. If I'm going to talk about the work we've done with traces, that will be a different session on its own. To give you an idea, a trace gives you a path for a particular request and response from the user to all the services that touched it. You can trace that path. The challenge with tracing is, unless you know exactly what trace you're looking for, it is super hard to use it to debug an issue accurately. It's not even a needle in a haystack. It's like looking for a leaf in a forest in the dark and you're blindfolded. It's hard.

## Ontology - Relationships, Not Just Connections

What we are really missing is not just connectedness, but the relationships between all these different pieces of this puzzle. We need to know the connections from these components, yes, but we also want to know how they interact with each other and how they are related to each other. This we call the observability ontology. Renzo will walk you through how we build these relationships and how we're putting this to use every day.

**Renzo Sanchez-Silva:** What is an ontology? It's just a formal specification of types, properties, and relationships. It's how we encode meaning. The basic construct of it is what you call the triple. The triple is a subject, a predicate, and an object, one fact per triple. It can be, a person is an employee. A service is important, or tier-0, for example. It can be properties too. This is the way we encode knowledge. It's very basic, very primordial, actually. It's like saying Tarzan loves Jane. Mommy loves Daddy. Babies have mini ontologies in their head. We just start collecting all these symbolic triples in our heads, and we start making propositions and reasoning and so forth. Ontologies are a way to encode knowledge. We're using W3C standards, RDF, OWL, web ontology language, SPARQL, the query language, SHACL is shapes constraint language, where we can say a service can be terminated in AWS, but then another constraint is if the service has more or less than 10% of capacity, you cannot terminate.

You can encode all these things. You can describe your services, the whole infrastructure. For example, we have API gateway. It's of type application. It's owned by a team. It is deployed in a region, and so forth. We can go and describe every single piece of the infrastructure, and every piece there is namespaced as well. We started building this baby ontology by describing incidents. We have identified a few namespaces. We're building an operational ontology. The goal, again, going back to Prasanna's vision, is to start finding this end-to-end observability graph of all the things connected in our infra. Let's call that the operational ontology. For instance, we identify these namespaces, and they can be referenced formally in the triples.

The problem is how to bring order to this operational chaos, in incidents scattered across Slack, PagerDuty, the MELT layer, metrics, traces, logs, and events. We all gather in the channel together. There are hundreds of messages, many services, people interacting. It's crazy until we find root cause, and then we fix the issue many hours later. What we want to do is observe that, start encoding in real time. Capturing the structures and how we preserve this in machine-readable types. A message can be, obviously, in English. Then someone says, API gateway has latency spiked after this deployment and so forth, and then Jane is rolling back. Then you can describe that in triples. That's the idea of doing this. Doing this by hand is hard. In fact, back in the day when there was no AI, you needed to have an ontologist who was always annotating, this is that, this is a person, and this is a service, and so forth.

With AI, we can do this much faster. For this, we created what we call a harvest pipeline. The incidents occur in Slack. That's where the action happens. For each channel, we transform every message, and we send every message to a bunch of enrichers. We call observability APIs. We use LLMs on top to find some extra classification, we correlate that. We save that in our graph database, we call it QuipuDB. The enrichers in green are very deterministic. We want to identify — for example, this is on incident.io, we use them as our incident management software — it's of this ID. We have the PagerDuty alert that was fired, for example, we extract that. We have various very deterministic, small algorithms that classify things because we know URLs, for example, we extract IDs. We extract, for example, also the sentiment of what is happening in the thread or the message.

In blue, we have alerts and the timeline that is occurring. People are sharing links to our dashboard. We encode that. We know what that means. Someone uploads a snapshot or a screenshot, we save that. It's a screenshot related to this graph, this metric. Then we use also LLMs at the end of this process where we have identified all these basic patterns or entities, and start finding new relationships or concepts, or new definitions. The LLM, you can see great things like this has canary analysis, deploy failed, and things like that. It starts finding concepts from what has been currently harvested. If it's new, we'll try to propose a new definition. This is what it's called in ontology, the open world assumption, where you are somewhat free to add relationships and we have ways to control that. It's a very controlled process. This is sequential. Sometimes if something fails, we just leave it. For example, we cannot grab some data, for example, you would leave it there. Then we check against our main metadata sources. For example, the software catalog is where we have the definition of all our services. We identify every service, and we know which type they are, and so forth. It's very controlled.

The goal is to create determinism from non-determinism. The non-determinism starts with the chaos in the channel, for example. That is the first source of non-determinism. Then we have a bunch of rules where we extract the data. Then, at the very end, the icing on the cake is the LLM coming and saying, this is probably a new relationship. There's a new connection here, and so forth. The paradox here is that we use stochastic tools to produce deterministic SPARQL queryable graph. The ontology is the bridge. It's where we save all the data. We identify five layers. This is a bootstrapping process. Deterministic extraction is those steps in green. We have a bunch of regexes, for example, that extract what a Jenkins job is by ID, by server. That's one entity. Those are very well-known patterns. Then the catalog validation is cataloging or checking against the service is a valid service deployed in this region and so forth.

We enrich with our APIs, observability APIs, and then we pass the LLM to classify and find new predicates. We also have a final step, what we call the probabilistic correlation, where we do the first search in all the conversation. We find chains where the incident started with this one application that affected application B, the original application was rolled back, and then the issue was fixed. That chain is identified also as an entity, a pattern that will be saved so that in the future, maybe we will find similarities, this incident with something in the future. Also, everything is constrained. We have categories, some taxonomies, and they are all constrained. When the LLM comes in, it wants to classify some category, say whether this conversation is an incident, question, discussion, or decision. It can only choose from those four values. We don't let it hallucinate new things unless maybe it tells us that, and then we'll see how we can manage that. Everything is typed through, for example, thread 123 is related, true is a Boolean.

The process we're describing is a bootstrapping process, what we call the knowledge flywheel, where every rotation, say a thread starts. We start identifying entities. More data comes in, we go around this flywheel by observing, enriching, inferring, and adapting. We observe the MELT data. We enrich with all these enrichers. We infer possibly new patterns or old patterns just to apply them, and we adapt to see what knowledge we can gather. The causal chain from one incident becomes a known pattern for the next. Observe, enrich, infer, that's the main process. In every iteration, we encode knowledge. Each cycle teaches the system, how the ontology will evolve this schema into memory. We have a process also where data is converting information, that information is encoded in knowledge. The fourth is wisdom, wisdom in the sense that we can find root cause. We can encode that today, but also in the future.

We're not predicting here yet. We're just describing always. Describing first. That's what a baby does. Daddy loves mommy, he knows that is a fact, but then more things will happen later. Everything is queryable via SPARQL. SPARQL is a weird syntax, but you get used to it. LLMs love it. We don't care. I don't even read it sometimes. Again, the ontology accumulates the structure of what keeps happening. Patterns become queryable facts.

It's not surprising that we use AI with this. We use Claude as our co-developer. When we get a new thread or a new message in an incident, we run this harvester. It creates a git worktree because we don't want to conflict with other branches, because there are many incidents happening, or maybe there's another iteration coming, the previous harvest can be working still on something and then it doesn't finish yet. It creates a git worktree or a branch. Then we use Claude also to harvest. We are not just running a script. We are running Claude with the harvesting script to extract the data because there's the LLM step that finds new patterns. At the same time, we use it to find maybe bugs in this very same script, so that it actually fixes itself. Then, say, that whole loop ends, and then we ended up with 15 branches only for this series of iterations, for one incident, and then human in the loop comes in.

Either me or Prasanna can come in and say, we got all these branches, and then we ask Claude to cherry pick and see what happened. It finds very interesting things. It finds new facts. Some, it says, these five branches did almost the same. They found this pattern. This other branch realized it was running on Sonnet 4.5, and it decided to upgrade itself to 4.6, which is cool. It's self-healing at the same time. You say, ok, merge. Then you keep going. It's a meta loop. The flywheel produces the knowledge, and the second one improves the machinery, and uses us for now. The idea is actually to make that non-dependent on humans and merge by itself. Then we continue iterating on top, so many flywheels on top, kind of a pyramid. Until it's completely autonomous, and then we move on to other domains or namespaces. This is how we see that in Slack, because Slack is convenient.

This is how it looks. For example, at each iteration, we find a few channels, and then, this happened to be a platform operations channel, a metrics/analytics engineering, all these classifications are something that it has found before. I never told it what it was, but I accepted it. I looked and spot checked. This is an alert management channel. It's not incident channels. These are general channels. We're also watching all the Slack channels as well. Because, remember, the idea is not only start with a small ontology and then go up into bigger parts of the infrastructure to describe them. This is just for general channels. This is what happens at the end of one iteration. For example, in this run it found some Turtle. Turtle is the way we serialize RDF in ontology language. It found some issues with the regex patterns, and the URIs were not saving correctly to valid Turtle.

It found a diff and so forth. That's what I can continue and say, you can merge this, and it goes and does it. All tests passed. The merge is complete and successful. Let me create a summary. That means for the next iteration, this code is already in the main branch. The next iteration, someone is adding a new message in the channel or a new discovery. Someone found a root cause that will actually run this new code. The next git worktree that spins up, it will already merge this fix.

The part that extracts these chains that I talked about, it's part of our correlation engine. It finds all these chains. For example, commit happened with retry logic. It was deployed with this version. Three minutes later there was an error, connection timeout, an alert fired, and the incident was started. Then the LLM comes up with its own confidence, but we save this chain. Many chains can happen in one incident. Imagine all these chains together, we'll have a lot of data in the future for fine-tuning and similarity. Agents are actually entropy reducers. Because, like I said, we receive noisy input. We spit out triples at the end, with new relationships, but at the same time, we find better ways of extracting data, because the agentic loop is happening. We do things, for example, like SME, Subject Matter Expert detection. This person is actually a Subject Matter Expert, and we find the topic too.

The accumulation of those triples will tell us eventually, this person is an expert in the service systems, or in Kafka and so forth. We also have these service topologies on a Turtle that already exists, because we build it from the ground up from many sources, using eBPF and logs and tracing. We created this graph of all our thousands of dependencies of services, which service calls each other. That's part of it. The contract is for the enrichers, we have a class, which is an interface where you can see the enrichment part, you can add triples to the graph and so forth. We have many of these, and more enrichers come as more data come in. Claude will actually tell us we need to build a new enricher because we're seeing this new thing that I don't know about from previous history and so forth. It's not in the code. It's not in the ontology. We need to create. Every URI in the entities are stable. We have a domain and IDs. We can always reference them by namespace.

From a graph, we go into a knowledge base. The process again, at every run of the harvester, we save the Turtle files in S3. We create a Parquet file. We have our graph database that is constantly polling that stream. It's actually a CDC, change data capture stream that loads those triples in memory. We can query in real time, because then we have the data. We have described the incident, for example. Someone else from a Claude Code session, or Cursor, or whatever can be querying what's going on with incident 123. We already have that data. The LLM will go make SPARQL queries and start asking questions, and then maybe correlate it to another incident in the past, and so forth. You use other MCP tools and so forth. We also save Blobs. For example, when someone uploads a graph, that file will be saved in S3, and we know it's related to a metric.

This is the query used. All that is described. The actual graph, the graph I mean the dashboard graph, it has a screenshot. The screenshot is saved, the bytes in S3. That Blob is also an entity. Everything is linked. We can recover all that data. Maybe you want to do forensic analysis and bring it back and get what happened exactly when things occurred. The ontology is a contract between chaos and understanding. Our stochastic engines are constrained to this very strong type system grounded on the ontology, and they emit deterministic triples that can be queryable. That's how we started with the incident ontology, a few namespaces, many agents and signals.

This is one of the artifacts that we produce on every run. This is just an HTML file saved in S3. We use this to debug, because I was just looking, is it extracting the right things? I know there has to be a Jira issue, or Slack users or file attachments and all those things. Every run, sometimes I realize, it's not finding file attachments. I need to have some artifacts, some goal, because I'm looking at the code all the time. I'm just looking at these results because I know what has to be in there. If it's not there, we need to keep fixing. Finally, by using this artifact, I could inspect. This is just a timeline of the incident. Everything is typed and linked. This is a big incident with many relationships. The topology, the service dependency is also encoded in there. Anyone can ask any questions, which teams are involved.

Who has low sentiment here? Who's very stressed, and all that? Everything is linked. Imagine all these incidents gluing up in the big graph so we can query them at the same time. The challenge is actually to keep all this data flowing in time, what we call the time travel. For now, we're just saving a few weeks of data. All this data is in S3, so eventually we can come up with a really big database that we can query at any time.

## AutoSRE - Putting the Ontology to Use

How do we use the ontology? One of the applications is what we call the AutoSRE. It's our AI SRE solution. That's the name now in the industry, for site reliability, aiding SRE with AI. We call it that. What AutoSRE does is taps into the metric, logs, and traces in real time, in also Slack. It has the topology, it has devices as well, and ways of looking at the end-to-end graph, which is the devices. Taps into the ontology, so in the knowledge graph, which is higher creative relationships as domain concepts. Going back to the LOLOMO incident or issue, AutoSRE comes in, not only the on-call, and it analyzes all the MELT data and versions of all the A/B tests that's running, and then it does correlation much faster than humans. It finds what incidents or apps are involved in it, and so forth. It finds logs as well.

It uses that. We have a bunch of MCPs and CLI tools querying the infrastructure, not only in ontology, but also really the data in general. It also looks at code, as many solutions out there do, because you cannot understand systems if you don't go down into the code. This is what it produces. It says, there was a GraphQL error spike for this record operation, and the next steps are these. We have actually added interactivity to this, where users can ask follow-up questions, because it's really hard to get any AI today to give you the right answer on the first go, in the first loop. You need to go through many iterations to dig in, because the LLMs can wander and stuff.

## What the Future Holds

What the future holds for us? There are three things. We want to find automatic root cause analysis. We don't want to keep digging. We can just give the crystal ball, this is the root cause. Next is to actually do auto-remediations. For example, if we know these patterns occur, the ontology, we'll say, in this situation, the best thing to do is just to terminate the instance. Fine. Do it. Auto-remediation. Without human interaction, even. The third one is self-healing infrastructure. That's the dream. We don't even want to get to an incident. We don't even want to get an alert. We want to keep shifting left and left, all the way to the point that the infrastructure is healing itself, because by the time we describe not only our incidents or our operational data, we will merge ontologies all the way to the data lineage to the point where, say, a movie is created.

There was a data error in a pipeline that was run by this team, by this person at this time, that cascaded into finally bad clicks on the user service. That's the end-to-end dream that we want to create. Not only forming the ontology, but merging into other bigger domains and create this big enterprise ontology that describes everything. That's going back to the original end-to-end wheel that Prasanna showed. Finding this huge cycle of connecting all the smaller ontologies into big.

Does anyone know this character from "Dark?" They are actually time travelers. There is a cabal that were able to harness time. They were able to go in the future, and actually, also knowledge. The inspiration here is I'm using one of their mantras. It says, "Der Anfang ist das Ende und das Ende ist der Anfang." "The beginning is the end and the end is the beginning." That means observe, enrich, infer. I just wanted to end that with that analogy to our show.

## Questions and Answers

**Participant 1:** How do you deal with deduplication and disambiguation in the graph on those incidents?

**Renzo Sanchez-Silva:** Right now, we're letting the LLM go into this open world. It does lookup, and it will create the triples, and say, no, this relationship is already there, and we let it create one. It is going to be a branch that has to be merged, but I'm just letting it pass, and it just keeps appending to the ontology. We may get to many similar concepts, maybe. We haven't checked that, but ideally, for example, for this smaller ontology, the process has stopped in the sense that it's becoming stable. I see less PRs that we need to merge, and so forth. It does eventually become stable. It does have to go through a human. Say an ontologist needs to actually see this, look at it closely. What I see is that deduplication is very controlled. This is possible.

**Prasanna Vijayanathan:** In the future, we also see the need to have a pattern extraction layer that can generalize patterns to help with deduplication. We're not there yet. We're able to manage.

**Participant 2:** In these graphs, everything can start relating to everything else. How do you fine-tune your infra to give you more signal over the noise, so that the things that it outputs are actually not something that has been hallucinated or possibly related, that it's actually actionable, real, and a signal for the user.

**Anand**: How do you remove false correlations, is that possible?

**Prasanna Vijayanathan:** Going back to old testing systems, we have a Swiss cheese model. We want to be able to give as much accurate and actionable information as possible, and the best person to tell us that would be the users right now. Like Renzo mentioned, it creates and self-heals itself, so that also exists. We have different quality metrics that we track. We do online logs, and then extract metrics from that, and then we have evals systems. It's like, we freeze our inputs and outputs, and then keep testing how our model is doing. Then we have human-in-the-loop feedback. All of those systems feed back into our model, and we haven't done that extensively yet, but it'll get into some kind of reinforcement learning to use those signals, and improve itself.

**Participant 3:** I'm a bit intrigued about the scale piece, because it seems like the whole graph is in Turtle, created by PRs like code. This is a small subset, then, of those 32 million per second queries, and we're not relating those bits. It's a much higher-level construct. Is that correct?

**Renzo Sanchez-Silva:** The idea is, for this high-volume throughput, we're just going to be collecting metadata, fingerprints of something has started. We have found this pattern, but we are not going to get the data on Turtle. There's no way to do that. We're going to find the fingerprint of what's happening, and then relate that for the rest of the graph. It's just always controlling cardinality as well.

**Participant 4:** How do you detect that something important didn't happen, instead of just picking up something bad that happened?

**Renzo Sanchez-Silva:** The ontologies actually have a very powerful formal reasoning. It's not AI-related. This is actually proportional, logical. Very hard algorithms that actually can reason, can find something that is incongruent or transitive or contradictory. We also plan to use formal reasoning to catch these things as things progress. Not only to find inconsistency in data, but also to find new connections. We can do formal reasoning with non-AI, but we can also use AI to help that base reasoning.

**Prasanna Vijayanathan:** A human in the loop also helps a lot in that system, where a developer notices something that was not caught, or we notice an incident or an issue that came up that actually had to be solved. Then we would look at why the system didn't catch it. In fact, recently we had a case where someone flagged it and said, why did you not tell me this? Then it was able to correct itself, and then create a new pair. That's the other side of it.

**Participant 5:** Could you tell me how often you would allow your ontology to change? How does that affect your knowledge graph rebuild?

**Renzo Sanchez-Silva:** We are letting it go with the open world assumption. It seems stable, at least for the incidents, but we're also going to find connections with alerts. That's a very small ontology, because all we want to catch is the metadata of the alerts. There are very few keywords or concepts in the alert, for example, and the signals. Then for other ones, like pipelines and jobs and all that, there's also a way to control that. It's very small. What is harder is the incidents, because anything can happen. The concepts are growing, but like I said, it becomes stable. At some point, you need to go back and look, we need to compress that with humans, and then it will stay controlled. We need human in the loop, for sure.

**Participant 6:** Any plans on open-sourcing?

**Prasanna Vijayanathan:** Maybe. Not yet.

**See more [presentations with transcripts](https://www.infoq.com/transcripts/presentations/)**
