How I Designed an AI Incident Response Agent with Hindsight A developer built an AI incident response agent that uses Hindsight as a persistent memory layer, letting the agent recall relevant past incidents alongside current evidence during an investigation and retain post-mortems for future use. The design keeps memory separate from the LLM's reasoning, with a small client wrapping the official hindsight-client SDK to expose retain and recall operations. The agent's flow runs from incident processing through AI investigation, recall of similar incidents, combined current and historical context, and finally retention of the post-mortem. How I Designed an AI Incident Response Agent with Hindsight When an incident happens, the alert usually isn't the whole story. There might be a similar incident from last week, a deployment that happened ten minutes earlier, or a previous post-mortem that already explains part of what we're seeing. If the agent only looks at the current alert, it misses that context. That was the problem I wanted to solve with my AI Incident Response Agent. I wanted the agent to investigate the current incident, but also have a way to look back at previous incidents when that history was actually useful. For that, I used Hindsight as the persistent memory layer. The architecture ended up being fairly simple. The interesting part was figuring out what to put into memory and how to bring the right memories back into an investigation. The basic architecture At a high level, the flow looks like this: Security Incident ↓ Incident Processing ↓ AI Investigation Agent ↓ Recall Relevant Incidents ↓ Current Evidence + Historical Context ↓ Investigation / Decision ↓ Retain Post-mortem There are really two directions here. The first is from the past into the current investigation: Previous incidents → Current investigation The second is from the current investigation into the future: Current investigation → Future incidents That second part is easy to overlook. Once an incident is finished, the investigation shouldn't necessarily disappear. Some of what we learned from it can become useful the next time something similar happens. That's where Hindsight fits. Why I didn't want a stateless agent A normal LLM workflow can do a good job if you give it all the information it needs in the prompt. But incident response doesn't really work that way. Imagine an alert like this: Service: payments-api Symptoms: Deployment: v2.4.1 deployed 12 minutes ago There is already quite a bit for the agent to investigate. But now imagine that the same service had a similar incident a few weeks ago. During that investigation, the team discovered that a particular deployment change caused almost the same symptoms. That information could be extremely useful. A stateless agent doesn't know about it unless we explicitly put it into the current context. That's what I wanted to change. Instead of making the agent start from zero every time, I wanted it to be able to ask: "Have we seen something like this before?" That sounds simple, but it changes the architecture quite a bit. Keeping memory separate from reasoning One decision I made was to keep the memory layer separate from the LLM itself. The LLM is responsible for reasoning about the incident. The application is responsible for orchestrating the investigation. Hindsight handles persistent memory. So instead of putting memory logic everywhere in the application, I created a small client around it. The rest of the incident-response code can then think in terms of: retain incident recall incidents rather than worrying about the details of the memory backend. That separation became useful because it keeps the responsibilities pretty clear. Turning an incident into memory The first half of the problem is retention. When an incident has been investigated, I want to preserve the useful parts of that investigation. Here's the core of the Hindsight integration: class HindsightMemoryClient: """Production wrapper around official hindsight-client SDK.""" php def retain incident self, incident: Incident - RetainResponse: """Retain a post-mortem into Hindsight with rich context and metadata.""" content = format incident as memory incident doc id = f"incident {incident.incident id}" return self.client.retain bank id=self.bank id, content=content, document id=doc id, metadata={ "incident id": incident.incident id, "service": incident.service, "severity": str incident.severity , "root cause": incident.root cause or "", "runbook": incident.runbook or "", }, tags= f"service:{incident.service}", f"severity:{incident.severity}", f"incident:{incident.incident id}", "type:incident", , There are a few things I care about here. First, I'm not just throwing the raw incident object into memory. format incident as memory turns the incident into something that can actually make sense as a stored investigation. I'm also giving every incident a predictable document ID: doc id = f"incident {incident.incident id}" Then I attach metadata such as the service, severity, root cause, and runbook. That matters because two incidents can have similar symptoms but completely different operational contexts. A critical incident in payments-api isn't necessarily comparable to a low-severity issue in an internal service just because both contain an authentication error. The tags give me another layer of structure: service:payments-api severity:critical incident:1234 type:incident So the memory isn't just a blob of text. It carries some of the context that explains what that memory actually represents. Finding the right memories Storing incidents is only half the problem. The more interesting part is retrieval. When a new incident arrives, I don't want to dump the entire incident history into the LLM. That would be noisy and expensive, and most of it probably wouldn't be relevant anyway. Instead, I build a query from the current incident. Here's the recall side: def recall incidents self, incident: Incident - list RecalledIncidentMemory : """Synthesize current symptoms, error logs, and deployment deltas into a semantic query.""" query str = f"Service: {incident.service}. " f"Symptoms: {'; '.join incident.symptoms }. " f"Errors: {' '.join incident.error logs :2 }. " f"Deployment: {incident.deployment version} {incident.deployment minutes before}m ago ." response = self.client.recall bank id=self.bank id, query=query str, max tokens=4096 I like this approach because the query represents the incident as a whole. It's not just: payments-api It's closer to: payments-api + current symptoms + recent errors + deployment information That gives the recall operation something meaningful to work with. The agent can then get historical incidents that actually have some relationship to the problem it is currently investigating. Keeping the recalled data useful Once Hindsight returns results, I map them into an application-level object: return RecalledIncidentMemory id=item.id, document id=item.document id, content=item.text, score=item.scores.get "total" if item.scores else None, tags=item.tags or , root cause=item.metadata.get "root cause" , resolution=item.metadata.get "resolution" , for item in response.results One detail I deliberately kept here is that I use the score returned by the memory system when it exists. I don't try to manufacture my own similarity score or make the result look more precise than it is. The application gets the recalled content, metadata, tags, and available score and can then pass the relevant information into the investigation flow. That keeps the memory layer fairly self-contained. What this looks like during an incident Let's take the payments-api example again. A deployment happens. Twelve minutes later, the service starts showing authentication failures and an elevated error rate. The current incident contains: Service: payments-api Symptoms: Deployment: v2.4.1 12 minutes ago The agent can investigate that information normally. But before reaching a conclusion, it can also recall previous incidents with similar characteristics. Suppose it finds an earlier incident involving the same service where authentication failures started shortly after a deployment. That doesn't mean the new incident has the same root cause. That's important. The previous incident is evidence worth considering, not an answer. The agent still has to look at the current deployment, current logs, and current symptoms. The useful context becomes: Current evidence + Relevant historical incidents ↓ Investigation That's the behavior I wanted from the memory layer. What happens when memory doesn't help? This was another thing I wanted to account for. Not every incident is going to have a useful historical match. Sometimes the system will be dealing with something completely new. Sometimes the previous incidents will look similar but turn out to have different causes. Sometimes there simply won't be enough historical information. That's fine. The agent shouldn't become dependent on memory to function. The fallback is straightforward: Relevant memory found ↓ Current evidence + historical context No relevant memory ↓ Current evidence The current incident always remains the primary source of evidence. Hindsight adds context when that context exists. The part I find most important The interesting thing about this architecture isn't that the agent can search old incidents. It's that completed investigations can become useful inputs for future investigations. The loop looks like this: Investigate ↓ Learn ↓ Retain ↓ Recall later ↓ Investigate with more context ↓ Learn again That gives the agent a kind of continuity. A new incident doesn't necessarily have to be the first time the system has encountered a particular pattern. What I learned building it It is easy to say that an agent should "remember everything." In practice, that's not what I want. The useful question is: What information could actually help with a future investigation? That changes what gets retained and how the memory is structured. A generic memory search isn't very useful if it returns unrelated incidents. Building the recall query from the current service, symptoms, errors, and deployment context gives historical retrieval a much stronger connection to the investigation. This is probably the most important rule in the system. A previous incident can suggest a direction for investigation. It shouldn't determine the conclusion. The current evidence still matters. The HindsightMemoryClient gives the rest of the application a simple interface for retaining and recalling incidents. That keeps the Hindsight-specific details in one place and makes the rest of the system easier to work with. A stateless agent answers the question in front of it. A memory-enabled agent can also use what it learned previously. For incident response, that distinction is useful because investigations naturally produce knowledge that can matter again later. Final Thoughts The biggest architectural decision in my AI Incident Response Agent wasn't the LLM itself. It was deciding what happens to the knowledge produced after an investigation is finished. If that knowledge disappears, the next investigation starts from scratch. If useful parts of it are retained, the next investigation has another source of context. That's why Hindsight became an important part of the architecture. The final loop is simple: Incident ↓ Investigate ↓ Retain what was learned ↓ New incident ↓ Recall relevant history ↓ Investigate with current + historical context The agent doesn't need to remember everything. It needs to remember the right things, retrieve them when they are relevant, and still validate them against what is happening now. That's the part of the system I'm most interested in: turning previous incident investigations from something that simply gets archived into something that can actively contribute to the next investigation. For the memory layer, I used Hindsight on GitHub and the Hindsight documentation. The broader idea of agent memory from Vectorize is also useful for understanding why persistent memory is different from simply putting more text into an LLM prompt. The agent doesn't have to start from zero every time.